← All IntelIntel / article 
AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin
- Source
- arxiv.org
- Author
- Bang Xie, Hao Liu, Zhenyu Shi, Yonghao Zhang, Senjian Zhang, Zhiyuan Peng, Xin Yin, Chenhao Ying, Yuan Luo, Wei Chen, Haiming Jin, Shaocong Long, Xu Liu, Zhe Peng
- Date

Why it matters
Repair agents that look competent on host run tests may not survive a real device build.
Terms in this piece · Glossary
- benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Read the source arxiv.org
Recommended reads
articleTowards the Generalizability of Leveraging ChatGPT in APR via Self-enhancing: An Empirical StudyQingyuan Li, Chuanyi Li, Yaopeng Yang, Ziwen Ge, Jidong Ge, Bin Luo
articleEvaluating Agentic Code Repair Capabilities in Distributed SystemsYibo Yan, Huijuan Wang, Junzhou He, Yizhuo Liang, Shaoyu Wang, Huanchen Sun, Seo Jin Park
articleRefine After Generation: Toward Correct and Concise Patches in LLM-based Program RepairWenqiang Luo, Jacky Keung, Xiaoyu Shi, Yicheng Sun, Boyang Yang, Zhou Yang, Haoye Tian
Comments
Checking sign-in…
Loading comments…