
Instead of retraining weights, this distills verified rewards each round into persistent memory files and repo state a fresh model session reads back, tested on IMO 2026, data science, and cybersecurity tasks.
articleAutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model ResearchMarjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
articleSelf-Evolving Embodied Agents via Skill-Harness EvolutionPeidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
articleFresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent MemoryEvan Chen, Shiqiang Wang, Christopher G. BrintonChecking sign-in…
Loading comments…