Vibeleaderboard
Index / tool
Visit evalplus.github.io
Category
Developer Tools
Rank
No. 3040Tools index

Previous survey · No. 2832 ·

Listed in
#8 Find AI benchmarks
Pricing
Open Source
Type
TOOL
Use case
Model & Agent Evaluation
Date

About

An evaluation framework that rebuilds HumanEval and MBPP with far more rigorous test suites, published as HumanEval+ and MBPP+. The extra tests exist to catch code that passes the originals' thin assertions but breaks on edge cases, and MBPP+ additionally narrows to 399 hand-verified tasks. Its leaderboard ranks models by pass@1 with greedy decoding and shows base and plus scores side by side, so the gap between them is the finding.

Why it made the leaderboard

It reruns HumanEval and MBPP with far more rigorous tests, catching generated code that passes the originals' thin assertions and fails on edge cases. Because the leaderboard prints base and plus scores side by side, the drop between them tells you how much of a published pass rate was test weakness rather than capability.

Tags

benchmarkhumanevalmbppcode-generationevaluationleaderboard

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.