BigCodeBench
bigcode-bench.github.io- Category
- Developer Tools
- Rank
- No. 1357Tools index
- Pricing
- Free
- Type
- TOOL
- Added
- Aug 19, 2026
About
A programming benchmark from the BigCode project that scores models on 1,140 tasks requiring compositional use of many libraries and function calls, rather than the self-contained puzzles earlier code benchmarks used. It runs in two modes: Complete, which tests completion from a docstring, and Instruct, which supplies only a natural-language instruction and tests whether the model turns intent into working code. Models rank by pass@1 with greedy decoding, and a roughly 150-task hard subset targets tougher real-world scenarios.
Why it made the leaderboard
It scores models on 1,140 tasks that require composing calls across many libraries, rather than the self-contained puzzles HumanEval popularised — closer to the code people actually write. The Instruct split gives only a natural-language instruction, so it measures turning intent into working code rather than completing a docstring.
Tags
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.