BigCodeBench
bigcode-bench.github.io- Category
- Developer Tools
- Rank
- No. 2848Tools index
Previous survey · No. 2683 ·
- Listed in
- #7 Find AI benchmarks
- Pricing
- Free
- Type
- TOOL
- Use case
- Model & Agent Evaluation
- Date
About
A programming benchmark from the BigCode project that scores models on 1,140 tasks requiring compositional use of many libraries and function calls, rather than the self-contained puzzles earlier code benchmarks used. It runs in two modes: Complete, which tests completion from a docstring, and Instruct, which supplies only a natural-language instruction and tests whether the model turns intent into working code. Models rank by pass@1 with greedy decoding, and a roughly 150-task hard subset targets tougher real-world scenarios.
Why it made the leaderboard
It scores models on 1,140 tasks that require composing calls across many libraries, rather than the self-contained puzzles HumanEval popularised — closer to the code people actually write. The Instruct split gives only a natural-language instruction, so it measures turning intent into working code rather than completing a docstring.
Tags
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.