Vibeleaderboard
← All Intel
Intel / article

Teaching Agents to Code Reliably

Source
arxiv.org
Author
Muhammad Ahmed Mohsin, Myeongsoo Kim, Kangrui Ruan, Shweta Garg, Varun Kumar, Murali Krishna Ramanathan
Date
Why it matters

Shows that diverse edit locations and verification against a reverted tree can raise coding-agent resolve rates while using about half the steps. Teams building coding-agent scaffolds can adopt both ideas.

Key takeaways · AI-distilled
  • The authors name three failure patterns: attempts keep returning to the same code location so extra samples add no coverage, different methods solve complementary issues no single run reaches, and an 's own tests accept many incorrect patches.
  • On 270 issues held out from training, weighted supervised raised pass@1 from 31.9% to 35.2% and pass@8 from 46.7% to 51.1%.
  • A reinforcement objective that trains the verifier on gold-labeled repairs and incorrect variants lifted pass@1 to 43.0% and pass@8 to 60.7%, raised verifier precision from 26.8% to 41.7%, and more than halved false acceptance.
  • Out of distribution, the authors report better resolution on two of three suites and better verifier precision on all three, with gains holding at 7B, 14B and 30B against published coder baselines.
Terms in this piece · Glossary
  • SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
  • fine-tuning — Taking a trained model and training it a bit more on your own examples so it gets better at one specific job.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Recommended reads
Comments

Checking sign-in…

Loading comments…