AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
- Source
- Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
- Author
- Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
- Date

Most coding- benchmarks score work against a known spec; this one measures whether an agent can make progress when the improvement direction is not given, which is closer to how research and open-ended engineering actually run.
“World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments.”
Marjan Moodi et al.
“This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks.”
Marjan Moodi et al.
“Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improve their starter on 63; in 91% of sessions the winning edit is a non-trivial research-style modification--a new objective, representation, rollout procedure, or architectural change--rather than a hyperparameter tweak.”
Marjan Moodi et al.
“Our benchmark offers a setting in which frontier coding agents can be evaluated on open-ended research rather than engineering-to-spec problems.”
Marjan Moodi et al.
articleGUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsLin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu, Yilun Zhao, Yu Rong
articleBelief-Calibrated Optimization: An Explicit World Model for Agentic OptimizationYuhan Chen, Zhihua Tian, Mahavir Dabas, Charith Peris, Rahul Gupta, Ming Jin, Feiyang Kang, Siyuan Zhang, Nan Wang, Ruoxi Jia
articleARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World ModelsZhengshu Zhang, Zhiyuan Li
Checking sign-in…
Loading comments…