
If the retry, not the feedback, is doing the work, self-repair loops are spending tokens for nothing.
“We argue that this comparison confounds the value of the feedback with the value of the extra attempt.”
“We attribute this to anchoring: when shown its previous attempt, a model reproduces a near-identical program in 33-68% of retries, against 2-14% under blind resampling.”
“Across six configurations spanning two families and two precisions, its magnitude is predicted by baseline quality alone (r=0.96) - the cost of anchoring is the cost of committing to a bad first attempt.”
postTraining models to explain their own failures, then rewarding the explanationWRITER
articleValidation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?Xiaonan Xu, Wenjing Wu
articleThe Recall Trap: A Recall-Maximizing Retriever Configuration Reduces Issue Resolution in Fixed-Budget Code ContextAlexander Adkins, Teimuraz TrapaidzeChecking sign-in…
Loading comments…