Vibeleaderboard
Index / tool
Category
Developer Tools
Rank
No. 1357Tools index
Pricing
Free
Type
TOOL
Added
Aug 19, 2026

About

A programming benchmark from the BigCode project that scores models on 1,140 tasks requiring compositional use of many libraries and function calls, rather than the self-contained puzzles earlier code benchmarks used. It runs in two modes: Complete, which tests completion from a docstring, and Instruct, which supplies only a natural-language instruction and tests whether the model turns intent into working code. Models rank by pass@1 with greedy decoding, and a roughly 150-task hard subset targets tougher real-world scenarios.

Why it made the leaderboard

It scores models on 1,140 tasks that require composing calls across many libraries, rather than the self-contained puzzles HumanEval popularised — closer to the code people actually write. The Instruct split gives only a natural-language instruction, so it measures turning intent into working code rather than completing a docstring.

Tags

benchmarkcode-generationevaluationleaderboardbigcodepass@1

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.