Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs
- Source
- AI Engineer
- Author
- AI Engineer
- Date
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
It gives engineers a real comparison of how current coding agents behave on legacy versus greenfield code, including the specific way a strong model fakes completion on a large refactor. The argument for judging reliability at 80–90% success rather than 50% is directly actionable when deciding how long to let an agent run unattended.
“it completed its goal in in 10 minutes and 22 seconds. And it only wrote 2,000 lines of code, which was a little bit fishy.”
Denys Linkov
“it's very easy to undergo AI psychosis, where you look at a deep research report that's 20 pages long”
Denys Linkov
“taking on technical debt and refactoring later is getting exponentially easier as the days go by”
Denys Linkov
“we can ship features that would take multiple months in under a week”
Denys Linkov
“sometimes it's good to pause, build a monorepo, and forge ahead”
Denys Linkov
videoI Gave an AI a Body — Cyrus Clarke, MIT Media Lab
videoThe Dark Arts of Skill Engineering — Paul Bakaus, Renaissance Geek
videoOperating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta
videoVertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
Checking sign-in…
Loading comments…


