
Quantifies that a passing test is not a proxy for a clean patch: diffs are consistently bloated, and minimality prompts do not fix it.
articleCausalRepair: Bridging the Causality Gap in Large Language Model-Based Automated Program Repair via Dual-SlicingLinhao Wu, Yizhou Chen, Zhen Yang, Pengyu Xue, Dan Hao
articleValidation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?Xiaonan Xu, Wenjing Wu
articleHow well LLM-based test generation techniques perform with newer LLM versions?Michael Konstantinou, Renzo Degiovanni, Mike PapadakisSign in to comment.
Loading comments…