Refine After Generation: Toward Correct and Concise Patches in LLM-based Program Repair
- Source
- Wenqiang Luo, Jacky Keung, Xiaoyu Shi, Yicheng Sun, Boyang Yang, Zhou Yang, Haoye Tian
- Author
- Wenqiang Luo, Jacky Keung, Xiaoyu Shi, Yicheng Sun, Boyang Yang, Zhou Yang, Haoye Tian
- Date

- SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Quantifies that a passing test is not a proxy for a clean patch: diffs are consistently bloated, and minimality prompts do not fix it.
“we find that even successful patches are consistently larger and more complex than developer patches, with the median approach producing 121.78% more total changes, 80.91% more net changes, and 43.99% higher cyclomatic complexity.”
Wenqiang Luo et al.
“In contrast, RECAP achieves a substantially better size-correctness tradeoff, cutting average total changes from +242.14% to +4.24% and net changes from +348.24% to -39.75% relative to developer patches while preserving or improving resolution by up to 42 instances.”
Wenqiang Luo et al.
“Across four host systems, prompting, commit-untangling, and minimality-aware baselines reduce patch size only by sacrificing 49 to 217 resolved instances.”
Wenqiang Luo et al.
“Our results indicate that minimality cannot be simply reduced to syntactic compression, and that decoupling minimization from generation offers a practical path to more reviewable repairs.”
Wenqiang Luo et al.
articleRethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiencyJunchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, Fabio Santos
articleTowards the Generalizability of Leveraging ChatGPT in APR via Self-enhancing: An Empirical StudyQingyuan Li, Chuanyi Li, Yaopeng Yang, Ziwen Ge, Jidong Ge, Bin Luo
articleAdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program RepairZ. C. Luo, J. C. Guo, W. J. He, S. Y. Wang, J. C. Yu, F. M. Zhao, Y. Chen, T. Cao, L. Q. Liu, N. Zheng, W. Xu, J. Jiang, Z. M. Zhao
Checking sign-in…
Loading comments…