
If you're building self-evolving agent skill systems, this shows gains are sparse and depend on including failed trajectories as feedback, and that more test-time compute (parallel sampling, sequential refinement) doesn't reliably substitute for persistent skill evolution.
“Evolution is sparse: only 55 of 388 candidates establish byte-distinct validation bests.”
“All 11 selections come from feedback conditions that include failed trajectories, although the relative ranking of Normal and Fail-only varies across settings.”
“Overall, persistent skill self-evolution is better understood as sparse, validation-filtered search with model- and benchmark-dependent returns, rather than steady improvement from additional rounds.”
articleSkillBoostHongqiang Lin, Chao Liu, Xiaofan Bai, Xuan Jin, Yuhong Li, Nenggan Zheng, Xipeng Cao
articleSelf-Evolving Embodied Agents via Skill-Harness EvolutionPeidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
post📄New Research on Self-Evolving Agents: When AI agents modify themselves, how…Tencent HyChecking sign-in…
Loading comments…