Zero2Repo: Can Coding Agents Build Repositories from Scratch?
Source
Pei Yang, Tianyu Shi, Yuhang Yao, Wanyi Chen, Tongyun Yang, Dun Pei, Haonan Wang, Pengbin Feng, Guanxu Yu, Jingchun Huang, Zeyu Zhang, Shuhan Sun, Hao Li, Xiang Li, Jie Xiao, Xinyu Wang, Hanxin Chen, Daqi Li, Qi Jia, Hongshan Lin, Zhizhou Gu, Zijun Tian, Weizhi Du, Lynn Ai, Eric Yang
Author
Pei Yang, Tianyu Shi, Yuhang Yao, Wanyi Chen, Tongyun Yang, Dun Pei, Haonan Wang, Pengbin Feng, Guanxu Yu, Jingchun Huang, Zeyu Zhang, Shuhan Sun, Hao Li, Xiang Li, Jie Xiao, Xinyu Wang, Hanxin Chen, Daqi Li, Qi Jia, Hongshan Lin, Zhizhou Gu, Zijun Tian, Weizhi Du, Lynn Ai, Eric Yang
Date
Key takeaways · AI-distilled
Each Zero2Repo task gives the AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → a product requirements document, an interface contract and an empty workspace. Tasks come from real, version-pinned open-source projects converted into specs, reproducible environments and hidden acceptance tests.
The hidden tests are validated both ways: a reference implementation derived from the upstream project must pass them, and adversarial validation must show they reject incorrect implementations. Scoring is binary, every test must pass, with no LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → judge.
On 11 tasks drawn from repositories frontier models have very likely seen in training, the strongest agent solved 10, and every failing submission still passed 90-99% of the hidden tests, the authors report.
For the two strongest agents, 67-100% of failed tests traced to a single omission or a low-frequency rule stated in the spec, not a missing subsystem, so the authors frame each failure as a concrete target for improvement.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Why it matters
Most coding-agent benchmarks test patching existing code. Zero2Repo measures whether agents can build a full repository from a requirements document and interface contract, with hidden tests validated against adversarial implementations.