
Most coding- benchmarks score work against a known spec; this one measures whether an agent can make progress when the improvement direction is not given, which is closer to how research and open-ended engineering actually run.
articleSelf-Evolving Embodied Agents via Skill-Harness EvolutionPeidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
articleRealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI FrameworksJinjun Huang, Zhongzhen Wen, Tongtong Xu, Meng Yan, Xin Xia, Zhongxin Liu
articleThe Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent BehaviorXiangzhe Xu, Hamidreza Saghir, Qianhui Wu, Marc-Alexandre C\^ot\'e, Tong Wang, Kiran Lakkaraju, Kexin Pei, Xiangyu ZhangSign in to comment.
Loading comments…