ZeroBench
zerobench.github.io- Category
- Developer Tools
- Rank
- No. 1294Tools index
- Listed in
- #41 Find AI benchmarks
- Pricing
- Free
- Platform
- web
- Type
- TOOL
- Date
About
Challenging visual reasoning questions with distinct main-question and subquestion scores. Compare pass@1, pass@5, and tool access separately.
What it does
A public leaderboard site tracking how large multimodal AI models perform on a fixed set of deliberately hard image interpretation puzzles, separating in house evaluations from scores that other labs and papers reported themselves, and flagging results where a model used external tools during testing.
Stated on the product site
- Question set
- The benchmark is built from a fixed set of 100 main questions plus 334 associated subquestions.
- Leaderboard structure
- Results are organized into three separate leaderboards: one for evaluations run by the benchmark's own team, one for scores that model developers or other labs reported themselves, and one preserving the original set of models tested at launch.
- Scoring metrics
- Each model's performance is reported using pass@1, pass@5, and pass^5 accuracy scores, along with average cost and token usage per question.
- Tool-assisted runs
- The site distinguishes model runs that used external tools, such as code execution or GUI interaction, from standard runs, marking the tool-assisted ones separately in its charts and tables.
- Error analysis
- The page also reports a categorization of common failure types behind incorrect model answers, derived from a multi-stage analysis process.
Not stated on the site
- It is not stated whether prospective evaluators can access the underlying question set and answer key directly, or only submit completed results for the maintainers to add.
- The page does not specify what licensing or usage terms govern the dataset or its images.
- It is not stated who funds or handles ongoing operation of the benchmark's public rankings beyond the individuals credited as paper authors.
Written from the product site at zerobench.github.io.
What it can do
Track official and externally reported model results across versions
Model version reports → Updated score records
Tags
benchmarkmultimodalllm-evaluationvision-languageleaderboardresearchgptclaude
Media

Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.