
Frontier mobile agents still fail on realistic multi-step phone workflows, and the ablations show how much retention choices, not just the model, drive that gap.
articleLegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy WorkflowsThilo Reintjes, Sivajeet Chand, Derui Zhu, Sushant Kumar Pandey, Alexander Pretschner
articleAutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model ResearchMarjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
articleAppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and KotlinBang Xie, Hao Liu, Zhenyu Shi, Yonghao Zhang, Senjian Zhang, Zhiyuan Peng, Xin Yin, Chenhao Ying, Yuan Luo, Wei Chen, Haiming Jin, Shaocong Long, Xu Liu, Zhe PengChecking sign-in…
Loading comments…