Vibeleaderboard
← All Intel
Intel / article

AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin

Source
arxiv.org
Author
Bang Xie, Hao Liu, Zhenyu Shi, Yonghao Zhang, Senjian Zhang, Zhiyuan Peng, Xin Yin, Chenhao Ying, Yuan Luo, Wei Chen, Haiming Jin, Shaocong Long, Xu Liu, Zhe Peng
Date
Why it matters

Repair agents that look competent on host run tests may not survive a real device build.

Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Recommended reads
Comments

Checking sign-in…

Loading comments…