E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch
Source
Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou
Author
Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou
Date
Key takeaways · AI-distilled
E2E-SWE has 186 whole-repository tasks across 11 languages. Given only a natural-language spec and an empty workspace, the AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → must build a complete, installable project that passes a comprehensive hidden test suite.
Each task is written by a software engineer working with an LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → to produce the tests and an implementation-independent spec, then audited by autonomous agents that repair defects using static inspection and failures seen in real model rollouts.
Across 13 frontier models, pass@1 ranged from 11.7% to 67.7%, which the authors say separates models clearly while leaving considerable headroom.
Trajectory analysis found long, front-loaded reasoning, which the authors read as a sign that building a working codebase from scratch depends on planning and system-level reasoning up front.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters
Measures whether coding agents can produce complete, installable repositories from a spec alone, across 11 languages, a more realistic test than localized patch benchmarks.