Vibeleaderboard

AI benchmark map

Choose benchmarks that test your task: repository coding, terminal work, reasoning, vision, security, or business workflows. Compare the benchmark version, agent harness, tool access, and scoring method before comparing model scores.

Surveyed 10 September 2026

Grouped by what they measure. Alphabetical within each area. Select a name for details.

Compare model results

Coding & software engineering

Write code, repair repositories, and work in a terminal.

Aider Polyglot Benchmark

Tests models on editing code across several programming languages through Aider. Scores depend on the model configuration and the editing format used.

View benchmark
ALE-Bench

Sakana AI's benchmark of long-horizon algorithm engineering on hard combinatorial optimization problems from programming contests. Scores are contest-style ratings, so compare the time budget and iteration setup.

View benchmark
AlgoTune

Tests whether models can write code that runs faster than expert reference implementations while remaining correct. Scores summarize speedups across many tasks, so read the aggregation method before comparing models.

View benchmark
BigCodeBench

Tests practical Python programming that combines multiple libraries. Generated solutions are checked against executable tests for tasks with detailed functional requirements.

View benchmark
Cline Bench

Evaluates agent performance on coding tasks through Cline. Read the task set and agent configuration to understand what a reported result represents.

View benchmark
Codeforces

Programming contests provide algorithmic problem-solving tasks for model evaluations. Reported ratings depend on the contest selection, sampling budget, and evaluation protocol.

View benchmark
CursorBench

Cursor's internal suite of ambiguous, multi-file coding tasks drawn from real Cursor sessions. Versions change the problem mix, so compare scores within one version, alongside cost and tokens per task.

View benchmark
DeepSWE

Evaluates coding agents on software engineering tasks. Check the evaluation version and agent configuration before comparing results across model releases.

View benchmark
EvalPlus

Tests generated code with expanded test suites for programming benchmarks. The additional cases help reveal incorrect solutions that smaller test sets can miss.

View benchmark
FrontierCode

Cognition's benchmark of hard, real open-source issues where coding agents must produce mergeable fixes. Epoch AI found too little public information to review it, so treat reported scores cautiously.

View benchmark
FrontierSWE

Ultra-long-horizon software engineering tasks spanning implementation, performance, scientific computing, visual reasoning, and AI research, with a twenty-hour budget per task. Compare the harness and budget before comparing agents.

View benchmark
GSO

Software optimization challenges where agents modify a program's code to make it substantially faster while staying correct. Compare the speedup threshold and evaluation harness alongside the pass rate.

View benchmark
HumanEval

Tests whether a model can complete small Python functions from prompts. Executable tests score correctness, with results affected by the number of sampled solutions.

View benchmark
LiveCodeBench

Tests code generation using programming contest problems collected over time. Results depend on the evaluation window and the model’s sampling configuration.

View benchmark
MirrorCode

Epoch AI's long-horizon coding benchmark. Models reimplement whole programs end to end from a specification without seeing the original source, so compare the task budget and harness alongside the score.

View benchmark
MLX Benchmark

Measures local model performance using Apple’s MLX framework. Hardware, quantization, model size, and runtime settings affect the reported speed and resource use.

View benchmark
NL2Repo-Bench

Tests whether agents can build complete Python repositories from natural-language specifications. Project test suites grade whether the generated code meets the requirements.

View benchmark
ProgramBench

Tests agents on building programs from specifications. Executable evaluation checks whether the resulting implementation satisfies the required behavior across the benchmark’s tasks.

View benchmark
ReactBench

Evaluates models on React development tasks. Use it to explore framework-specific coding performance, while checking how solutions are tested and graded.

View benchmark
SciCode

Scientist-curated research-coding problems across physics, chemistry, biology, and materials science. Each problem decomposes into subproblems, so check whether a reported score counts main problems or subproblems.

View benchmark
SWE-bench

Tests whether coding agents can resolve real GitHub issues. Results depend on the benchmark split, repository environment, and agent harness used.

View benchmark
SWE-bench Live

Tests coding agents on recent repository issues. The benchmark updates over time, so compare task dates and language coverage alongside reported scores.

View benchmark
SWE-rebench

Evaluates repository issue resolution with refreshed tasks. Its changing task pool helps track coding performance while reducing dependence on familiar public test cases.

View benchmark
Terminal-Bench

Tests whether agents can complete practical tasks in a terminal. Compare the benchmark version and execution environment alongside the reported success rate.

View benchmark
WeirdML

Tests whether models can write working PyTorch code for unfamiliar, well-specified machine-learning problems under a fixed compute budget. Scores come from a non-agentic setup, so agent results are not comparable.

View benchmark

Reasoning, science & mathematics

Solve hard questions and generalize to unfamiliar problems.

AIME

The American Invitational Mathematics Examination, a high-school olympiad qualifier with integer answers. Each year's paper is a separate edition; compare the same year, sample count, and sampling settings.

View benchmark
ARC-AGI

Tests generalization to unfamiliar tasks. Compare the specific ARC-AGI version and evaluation rules, because task formats and permitted interaction differ between releases.

View benchmark
ArXivMath

Tests research mathematics from recent arXiv papers. Compare monthly editions, native agent harnesses, tool access, and final-answer grading; August 2026 prioritizes results that resolve earlier conjectures.

View benchmark
BeyondAIME

ByteDance Seed's contamination-resistant set of competition problems at or above late-AIME difficulty, each rewritten so the integer answer is as hard to guess as to derive.

View benchmark
BrokenArXiv

Tests whether models recognize false mathematical conjectures instead of claiming to prove them. The August 2026 edition uses native agent harnesses and a revised 0–3 rubric, not binary accuracy.

View benchmark
FrontierMath

Epoch AI's held-out, expert-written mathematics problems, ranging from advanced undergraduate through research level across tiers. Compare the tier, version, and sampling budget before comparing model scores.

View benchmark
GPQA Diamond

Tests graduate-level reasoning in biology, chemistry, and physics. Diamond is an expert-validated subset; check the subset and sampling setup when comparing scores.

View benchmark
HMMT February

The Harvard-MIT Mathematics Tournament's February competition, with final-answer problems across algebra, combinatorics, geometry, and number theory. Each year is a separate edition; compare the same year and sample count.

View benchmark
Humanity's Last Exam

Tests models on difficult questions across expert domains. Full, text-only, and tool-assisted evaluations have different scopes and should be compared separately.

View benchmark
MathArena Apex

Tests mathematical reasoning with difficult competition problems. Compare the problem-set edition and sampling budget alongside the percentage of correct final answers.

View benchmark
SimpleQA Verified

One thousand filtered factoid questions measuring parametric recall without tools. Headline accuracy also reflects a model's willingness to guess, so read abstention and calibration alongside the score.

View benchmark
Terminal-Bench-Science

Expert-authored research workflows from the life, physical, earth, mathematical, and engineering sciences, completed by agents in a terminal. Releases change the task set; compare the same release, harness, and trial count.

View benchmark

Agents & workflows

Use tools and complete work across applications and environments.

Agents' Last Exam

Tests agents on extended professional workflows across multiple domains. Compare the shared evaluation framework, agent harness, and task configuration alongside reported results.

View benchmark
AutomationBench

Tests agents on business workflows across simulated applications. Grading checks the final system state for tasks spanning functions such as sales, support, and finance.

View benchmark
Berkeley Function Calling Leaderboard

Tests models on choosing and calling functions correctly. Evaluation categories distinguish tool selection, argument construction, and more complex interaction patterns.

View benchmark
DeepResearchBench

Tests whether agents can research questions on the live web and synthesize correct answers. Results depend on the browsing harness and the date of the run, so compare both alongside scores.

View benchmark
DevTool Arena

Compares development environments and sandboxes through practical tasks. Results describe the tested infrastructure and workflow, so they are not standalone model capability scores.

View benchmark
DrivingBench

Puts a frontier model in control of a real Toyota Corolla's steering, throttle, and brakes through three MCP tools and scores progress along a fixed cone course. Compare attempt counts and harness latency before comparing progress.

View benchmark
GAIA

Tests assistants on questions requiring reasoning, information retrieval, and tool use. Difficulty levels and available tools influence the reported completion rate.

View benchmark
GDPval

OpenAI's evaluation of well-specified workplace tasks drawn from real occupations across nine economic sectors. Grading uses expert comparison against human deliverables, so read the win-rate definition before comparing models.

View benchmark
MLE-bench

Tests agents on machine-learning engineering tasks drawn from competitions. Results reflect the full workflow, including data preparation, training, and submission quality.

View benchmark
OSWorld

Tests agents on tasks in desktop applications. Evaluation checks the resulting computer state, with performance affected by the model’s interaction setup and task budget.

View benchmark
PostTrainBench

Tests whether CLI agents can post-train small base models under a fixed compute budget, scored on seven target benchmarks. Epoch AI verified the design, but compare the budget and harness alongside results.

View benchmark
SkillsBench

Tests how reusable skills affect agent task performance. Compare the skill configuration, model, and harness to distinguish improvements from changes in the setup.

View benchmark
ToolBench

Evaluates models on solving tasks with external tools. Check the tool collection, evaluation method, and execution conditions before comparing reported results.

View benchmark
Vending-Bench 2

Tests whether an agent can run a simulated vending machine business profitably and stay coherent over a full simulated year. Compare the run length and tool setup before comparing final balances.

View benchmark
WebArena

Tests agents on tasks across realistic websites. Success is graded against the required outcome, covering workflows that involve navigation and changing application state.

View benchmark
τ²-bench

Tests agents in tool-using conversations with simulated users. Success depends on following domain policies and reaching the required state across the interaction.

View benchmark

Security

Find, reproduce, and assess software vulnerabilities.

Cybench

Cybersecurity capture-the-flag challenges measuring autonomous vulnerability discovery and exploitation in sandboxed environments. Compare the guidance setting and agent scaffold before comparing solve rates.

View benchmark
CyberGym

Tests whether agents can reproduce real software vulnerabilities. Agents receive a description and unpatched codebase, then produce executable proof-of-concept inputs.

View benchmark
ExploitBench

Measures how far agents can take real, hardened vulnerabilities toward a working exploit, scored on a ladder of verifiable capability tiers. Few exploits are public, which limits contamination but complicates reproduction.

View benchmark
ExploitGym

Tests whether agents can turn known vulnerabilities into working exploits. Evaluation runs in controlled environments and distinguishes exploitation from simply reproducing a crash.

View benchmark
SEC-Bench Pro

Tests long-horizon security bug hunting in complex systems. Evaluation uses reproducible validation for tasks involving targets such as browser engines and the Linux kernel.

View benchmark

Vision & design

Read images and charts, recognize objects, and judge visual output.

BabyVision

Tests visual discrimination, tracking, spatial perception, and pattern recognition. Direct vision answers and tool-assisted results use different setups and should be compared separately.

View benchmark
Chartography

Tests understanding of real professional charts using expert-written questions. Charts span domains including science, finance, healthcare, engineering, and manufacturing.

View benchmark
DesignArena

Compares visual output through arena-style evaluations. Preference results depend on the task and judging setup, rather than a universal measure of design quality.

View benchmark
Roboflow Model Leaderboard

Compares computer-vision models on published evaluation tasks. Check the dataset, model configuration, and task-specific metric to interpret accuracy and performance results.

View benchmark
ZeroBench

Tests difficult visual reasoning with separate main-question and subquestion scores. Compare sampling metrics and tool access before treating two reported results as equivalent.

View benchmark

Embeddings & retrieval

Represent meaning, find relevant content, and rank results.

MTEB

Evaluates embedding models across multiple language tasks. Compare the benchmark version, language coverage, and task subset, especially when choosing models for retrieval.

View benchmark
Reranker Simple Benchmark

Compares reranking models on ordering retrieved content. Check the dataset, candidate set, and ranking metric to judge relevance to your search workflow.

View benchmark

Speech & audio

Recognize speech and evaluate spoken interactions.

UltraEval-Audio

Evaluates audio models across supported speech and audio tasks. Results depend on the dataset and task-specific metric, so compare matching evaluation configurations.

View benchmark

Model & infrastructure comparisons

Compare broad capability, human preference, speed, and hardware fit.

Arena

Compares models using human preferences in paired interactions. Ratings reflect the voting population, category, and collection period, rather than objective correctness alone.

View benchmark
Artificial Analysis

Publishes model evaluations and performance comparisons. Read each metric’s methodology and tested configuration when comparing capability, speed, cost, or human preference.

View benchmark
Epoch AI Benchmarking Hub

Epoch AI's hub of independently run benchmark results, the Epoch Capabilities Index, and benchmark reviews rating evaluations Verified or Flawed. Read each review before trusting a benchmark's headline numbers.

View benchmark
InferenceX

Compares inference performance across model and hardware configurations. Throughput and latency depend on the workload, serving stack, and resource allocation used.

View benchmark
LiveBench

Evaluates models on regularly refreshed tasks across several capabilities. Compare the release date and task category because the question set changes over time.

View benchmark
METR Time Horizons

Estimates the length of software tasks a model completes correctly more often than not. Horizons are model estimates from a fixed task suite, so compare the suite version and confidence intervals.

View benchmark
Open Weights

Helps compare openly available models. Check the evidence behind each comparison, including the model version, evaluation source, and deployment configuration.

View benchmark
WhichLLM

Helps explore model choices through published comparisons. Verify the source and date of each result before using it to select a model for your workload.

View benchmark

Specialist tasks & behavior

Legal work, ethical judgment, uncertainty, games, and constrained training.

AI Battlegrounds

Evaluates agents through game-based competition. Outcomes depend on the game rules, opponents, and agent configuration rather than a general measure of model capability.

View benchmark
BullshitBench

Tests how models respond to flawed or nonsensical prompts. It examines whether a model challenges the premise instead of confidently inventing an answer.

View benchmark
Harvey LAB (Legal Agent Benchmark)

Evaluates agents on legal work. Read the task definitions and grading methodology to understand which parts of a legal workflow the results support.

View benchmark
KellyBench

Tests long-horizon decision-making in simulated sports betting markets. Agents build models, size bets, manage risk, and adapt their strategy across a football season.

View benchmark
Parameter Golf

Compares language-model training approaches under tight resource constraints. Results reflect the training rules and model-size budget as well as the final evaluation score.

View benchmark
Philosophy Bench

Evaluates model responses to philosophical questions. Interpretation depends on the prompts and judging methodology, rather than a single objective notion of philosophical correctness.

View benchmark
Vals Minecraft Research Run

Reports progress from one 141-hour GPT-6 Astra Minecraft run in Normal Survival with keepInventory enabled. These observations are not a standardized cross-model leaderboard or speed ranking.

View benchmark

What to look for

  • 01Does the benchmark test the task you care about? Multiple-choice reasoning tells you nothing about multi-file refactoring.
  • 02Is contamination plausible? Public benchmarks leak into training data and scores drift upward without capability doing the same.
  • 03How is success graded: tests, exact answers, human review, or a model judge? Read the grading rules and known limitations.

Common questions

Where are the benchmarks from the DeepSeek V4.1 Flash launch?
All 19 launch evaluations are represented here by 16 benchmark entries. Terminal-Bench 2.1, 3.0, and 4.0 share one entry; HLE with and without tools share Humanity's Last Exam. DeepSWE v1.1, GPQA Diamond, MathArena Apex, and ZeroBench-main retain their specific scope. The Model Atlas carries the reported scores, harness details, and source discrepancies.
What is SWE-bench?
A benchmark built from real GitHub issues and their merged fixes. The model gets the repository and the issue, and its patch is scored by whether the project's own tests pass.
Why does the top benchmark model not feel best in practice?
Benchmarks measure isolated task completion under ideal context. Daily work is about following your conventions, handling ambiguity, and staying useful over a long session — none of which is scored.
Is there a benchmark for my specific domain?
Increasingly, yes. Legal reasoning, ethical dilemmas, embeddings, object detection, speech and React code all have dedicated benchmarks now, and a narrow one measuring your actual task usually tells you more than a general leaderboard does.

More in Understand the AI landscape