
Comments copied from correct solutions raise coding-model pass@1 by 17.2%, but comments describing a wrong or unrelated problem cut it by up to 20.8%, and most models can't self-correct that gap, a concrete caution for scaffolding coding- prompts with comments.
“Comments from source solutions that pass the tests raise recipient pass@1 by 17.2% on average.”
“comments describing failed solutions provide no reliable gain, while comments written for a different problem reduce pass@1 by 20.8%”
“most recipient models show no significant recovery of the external-comment gain, and the best case recovers only 24%”
“These results show that comments help code generation not merely because they are comments, but because they can provide correct solution content that prompting cannot reliably elicit.”
articleCriticGen: Generation-Aware Evaluation as Actionable FeedbackHuifang Du, Zecheng Zuo, Sen Wang, Chenghao Fan, Haofen Wang, Yehui Yang
articleTIPCODER: Reinforcement Learning Boosted Test-time Instruction Proposer for Code GenerationMinyu Chen, Sihao Wu, Ling-I Wu, Song Qin, Jingyang Li, Lei Ning, Jianxin Xue, Guoqiang Li
articleCode Health in LLM-Based Test Generation: Effectiveness and Token EfficiencyFreya Wirdemann, Markus Borg, Nadim Hagatulah, Adam TornhillChecking sign-in…
Loading comments…