Vibeleaderboard
← All Intel
Intel / article

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

Source
Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou
Author
Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou
Date
Key takeaways · AI-distilled
  • E2E-SWE has 186 whole-repository tasks across 11 languages. Given only a natural-language spec and an empty workspace, the must build a complete, installable project that passes a comprehensive hidden test suite.
  • Each task is written by a software engineer working with an to produce the tests and an implementation-independent spec, then audited by autonomous agents that repair defects using static inspection and failures seen in real model rollouts.
  • Across 13 frontier models, pass@1 ranged from 11.7% to 67.7%, which the authors say separates models clearly while leaving considerable headroom.
  • Trajectory analysis found long, front-loaded reasoning, which the authors read as a sign that building a working codebase from scratch depends on planning and system-level reasoning up front.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters

Measures whether coding agents can produce complete, installable repositories from a spec alone, across 11 languages, a more realistic test than localized patch benchmarks.

Recommended reads
Comments

Checking sign-in…

Loading comments…