Vibeleaderboard
Index / tool
Visit zerobench.github.io
Category
Developer Tools
Rank
No. 1294Tools index
Listed in
#41 Find AI benchmarks
Pricing
Free
Platform
web
Type
TOOL
Date

About

Challenging visual reasoning questions with distinct main-question and subquestion scores. Compare pass@1, pass@5, and tool access separately.

What it does

A public leaderboard site tracking how large multimodal AI models perform on a fixed set of deliberately hard image interpretation puzzles, separating in house evaluations from scores that other labs and papers reported themselves, and flagging results where a model used external tools during testing.

Stated on the product site

Question set
The benchmark is built from a fixed set of 100 main questions plus 334 associated subquestions.
Leaderboard structure
Results are organized into three separate leaderboards: one for evaluations run by the benchmark's own team, one for scores that model developers or other labs reported themselves, and one preserving the original set of models tested at launch.
Scoring metrics
Each model's performance is reported using pass@1, pass@5, and pass^5 accuracy scores, along with average cost and token usage per question.
Tool-assisted runs
The site distinguishes model runs that used external tools, such as code execution or GUI interaction, from standard runs, marking the tool-assisted ones separately in its charts and tables.
Error analysis
The page also reports a categorization of common failure types behind incorrect model answers, derived from a multi-stage analysis process.

Not stated on the site

  • It is not stated whether prospective evaluators can access the underlying question set and answer key directly, or only submit completed results for the maintainers to add.
  • The page does not specify what licensing or usage terms govern the dataset or its images.
  • It is not stated who funds or handles ongoing operation of the benchmark's public rankings beyond the individuals credited as paper authors.

Written from the product site at zerobench.github.io.

What it can do

  • Track official and externally reported model results across versions

    Model version reportsUpdated score records

Tags

benchmarkmultimodalllm-evaluationvision-languageleaderboardresearchgptclaude

Media

ZeroBench

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.