Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Rank
No. 2206Tools index
Listed in
#14 Find AI benchmarks
Type
TOOL
GitHub
173 stars
Date

About

Tests whether coding agents can build complete Python repositories from natural-language specifications, scored against the projects’ test suites.

What it can do

  • Evaluate coding agents by generating a complete Python repository from a natural-language specification

    Natural-language specificationPython repository

  • Score generated repositories against the project's test suite

    Generated Python repository and test suiteTest results/score

Why it made the leaderboard

Compare the task, benchmark version, harness, and grading method before using model scores to choose a model.

Tags

benchmarkevaluation

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.