Vibeleaderboard
← All Intel
Intel / article

Refine After Generation: Toward Correct and Concise Patches in LLM-based Program Repair

Source
Wenqiang Luo, Jacky Keung, Xiaoyu Shi, Yicheng Sun, Boyang Yang, Zhou Yang, Haoye Tian
Author
Wenqiang Luo, Jacky Keung, Xiaoyu Shi, Yicheng Sun, Boyang Yang, Zhou Yang, Haoye Tian
Date
Terms in this piece · Glossary
  • SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters

Quantifies that a passing test is not a proxy for a clean patch: diffs are consistently bloated, and minimality prompts do not fix it.

Key quotes

“we find that even successful patches are consistently larger and more complex than developer patches, with the median approach producing 121.78% more total changes, 80.91% more net changes, and 43.99% higher cyclomatic complexity.”

Wenqiang Luo et al.

“In contrast, RECAP achieves a substantially better size-correctness tradeoff, cutting average total changes from +242.14% to +4.24% and net changes from +348.24% to -39.75% relative to developer patches while preserving or improving resolution by up to 42 instances.”

Wenqiang Luo et al.

“Across four host systems, prompting, commit-untangling, and minimality-aware baselines reduce patch size only by sacrificing 49 to 217 resolved instances.”

Wenqiang Luo et al.

“Our results indicate that minimality cannot be simply reduced to syntactic compression, and that decoupling minimization from generation offers a practical path to more reviewable repairs.”

Wenqiang Luo et al.
Recommended reads
Comments

Checking sign-in…

Loading comments…