Vibeleaderboard
Index / tool
Visit bigcode-bench.github.io
Category
Developer Tools
Rank
No. 2848Tools index

Previous survey · No. 2683 ·

Listed in
#7 Find AI benchmarks
Pricing
Free
Type
TOOL
Use case
Model & Agent Evaluation
Date

About

A programming benchmark from the BigCode project that scores models on 1,140 tasks requiring compositional use of many libraries and function calls, rather than the self-contained puzzles earlier code benchmarks used. It runs in two modes: Complete, which tests completion from a docstring, and Instruct, which supplies only a natural-language instruction and tests whether the model turns intent into working code. Models rank by pass@1 with greedy decoding, and a roughly 150-task hard subset targets tougher real-world scenarios.

Why it made the leaderboard

It scores models on 1,140 tasks that require composing calls across many libraries, rather than the self-contained puzzles HumanEval popularised — closer to the code people actually write. The Instruct split gives only a natural-language instruction, so it measures turning intent into working code rather than completing a docstring.

Tags

benchmarkcode-generationevaluationleaderboardbigcodepass@1

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.