AI benchmark map
Choose benchmarks that test your task: repository coding, terminal work, reasoning, vision, security, or business workflows. Compare the benchmark version, agent harness, tool access, and scoring method before comparing model scores.
Surveyed 10 September 2026
Grouped by what they measure. Alphabetical within each area. Select a name for details.
Compare model resultsCoding & software engineering
Write code, repair repositories, and work in a terminal.
Aider Polyglot Benchmark
Tests models on editing code across several programming languages through Aider. Scores depend on the model configuration and the editing format used.
View benchmarkALE-Bench
Sakana AI's benchmark of long-horizon algorithm engineering on hard combinatorial optimization problems from programming contests. Scores are contest-style ratings, so compare the time budget and iteration setup.
View benchmarkAlgoTune
Tests whether models can write code that runs faster than expert reference implementations while remaining correct. Scores summarize speedups across many tasks, so read the aggregation method before comparing models.
View benchmarkBigCodeBench
Tests practical Python programming that combines multiple libraries. Generated solutions are checked against executable tests for tasks with detailed functional requirements.
View benchmarkCline Bench
Evaluates agent performance on coding tasks through Cline. Read the task set and agent configuration to understand what a reported result represents.
View benchmarkCodeforces
Programming contests provide algorithmic problem-solving tasks for model evaluations. Reported ratings depend on the contest selection, sampling budget, and evaluation protocol.
View benchmarkCursorBench
Cursor's internal suite of ambiguous, multi-file coding tasks drawn from real Cursor sessions. Versions change the problem mix, so compare scores within one version, alongside cost and tokens per task.
View benchmarkDeepSWE
Evaluates coding agents on software engineering tasks. Check the evaluation version and agent configuration before comparing results across model releases.
View benchmarkEvalPlus
Tests generated code with expanded test suites for programming benchmarks. The additional cases help reveal incorrect solutions that smaller test sets can miss.
View benchmarkFrontierCode
Cognition's benchmark of hard, real open-source issues where coding agents must produce mergeable fixes. Epoch AI found too little public information to review it, so treat reported scores cautiously.
View benchmarkFrontierSWE
Ultra-long-horizon software engineering tasks spanning implementation, performance, scientific computing, visual reasoning, and AI research, with a twenty-hour budget per task. Compare the harness and budget before comparing agents.
View benchmarkGSO
Software optimization challenges where agents modify a program's code to make it substantially faster while staying correct. Compare the speedup threshold and evaluation harness alongside the pass rate.
View benchmarkHumanEval
Tests whether a model can complete small Python functions from prompts. Executable tests score correctness, with results affected by the number of sampled solutions.
View benchmarkLiveCodeBench
Tests code generation using programming contest problems collected over time. Results depend on the evaluation window and the model’s sampling configuration.
View benchmarkMirrorCode
Epoch AI's long-horizon coding benchmark. Models reimplement whole programs end to end from a specification without seeing the original source, so compare the task budget and harness alongside the score.
View benchmarkMLX Benchmark
Measures local model performance using Apple’s MLX framework. Hardware, quantization, model size, and runtime settings affect the reported speed and resource use.
View benchmarkNL2Repo-Bench
Tests whether agents can build complete Python repositories from natural-language specifications. Project test suites grade whether the generated code meets the requirements.
View benchmarkProgramBench
Tests agents on building programs from specifications. Executable evaluation checks whether the resulting implementation satisfies the required behavior across the benchmark’s tasks.
View benchmarkReactBench
Evaluates models on React development tasks. Use it to explore framework-specific coding performance, while checking how solutions are tested and graded.
View benchmarkSciCode
Scientist-curated research-coding problems across physics, chemistry, biology, and materials science. Each problem decomposes into subproblems, so check whether a reported score counts main problems or subproblems.
View benchmarkSWE-bench
Tests whether coding agents can resolve real GitHub issues. Results depend on the benchmark split, repository environment, and agent harness used.
View benchmarkSWE-bench Live
Tests coding agents on recent repository issues. The benchmark updates over time, so compare task dates and language coverage alongside reported scores.
View benchmarkSWE-rebench
Evaluates repository issue resolution with refreshed tasks. Its changing task pool helps track coding performance while reducing dependence on familiar public test cases.
View benchmarkTerminal-Bench
Tests whether agents can complete practical tasks in a terminal. Compare the benchmark version and execution environment alongside the reported success rate.
View benchmarkWeirdML
Tests whether models can write working PyTorch code for unfamiliar, well-specified machine-learning problems under a fixed compute budget. Scores come from a non-agentic setup, so agent results are not comparable.
View benchmarkReasoning, science & mathematics
Solve hard questions and generalize to unfamiliar problems.
AIME
The American Invitational Mathematics Examination, a high-school olympiad qualifier with integer answers. Each year's paper is a separate edition; compare the same year, sample count, and sampling settings.
View benchmarkARC-AGI
Tests generalization to unfamiliar tasks. Compare the specific ARC-AGI version and evaluation rules, because task formats and permitted interaction differ between releases.
View benchmarkArXivMath
Tests research mathematics from recent arXiv papers. Compare monthly editions, native agent harnesses, tool access, and final-answer grading; August 2026 prioritizes results that resolve earlier conjectures.
View benchmarkBeyondAIME
ByteDance Seed's contamination-resistant set of competition problems at or above late-AIME difficulty, each rewritten so the integer answer is as hard to guess as to derive.
View benchmarkBrokenArXiv
Tests whether models recognize false mathematical conjectures instead of claiming to prove them. The August 2026 edition uses native agent harnesses and a revised 0–3 rubric, not binary accuracy.
View benchmarkFrontierMath
Epoch AI's held-out, expert-written mathematics problems, ranging from advanced undergraduate through research level across tiers. Compare the tier, version, and sampling budget before comparing model scores.
View benchmarkGPQA Diamond
Tests graduate-level reasoning in biology, chemistry, and physics. Diamond is an expert-validated subset; check the subset and sampling setup when comparing scores.
View benchmarkHMMT February
The Harvard-MIT Mathematics Tournament's February competition, with final-answer problems across algebra, combinatorics, geometry, and number theory. Each year is a separate edition; compare the same year and sample count.
View benchmarkHumanity's Last Exam
Tests models on difficult questions across expert domains. Full, text-only, and tool-assisted evaluations have different scopes and should be compared separately.
View benchmarkMathArena Apex
Tests mathematical reasoning with difficult competition problems. Compare the problem-set edition and sampling budget alongside the percentage of correct final answers.
View benchmarkSimpleQA Verified
One thousand filtered factoid questions measuring parametric recall without tools. Headline accuracy also reflects a model's willingness to guess, so read abstention and calibration alongside the score.
View benchmarkTerminal-Bench-Science
Expert-authored research workflows from the life, physical, earth, mathematical, and engineering sciences, completed by agents in a terminal. Releases change the task set; compare the same release, harness, and trial count.
View benchmarkAgents & workflows
Use tools and complete work across applications and environments.
Agents' Last Exam
Tests agents on extended professional workflows across multiple domains. Compare the shared evaluation framework, agent harness, and task configuration alongside reported results.
View benchmarkAutomationBench
Tests agents on business workflows across simulated applications. Grading checks the final system state for tasks spanning functions such as sales, support, and finance.
View benchmarkBerkeley Function Calling Leaderboard
Tests models on choosing and calling functions correctly. Evaluation categories distinguish tool selection, argument construction, and more complex interaction patterns.
View benchmarkDeepResearchBench
Tests whether agents can research questions on the live web and synthesize correct answers. Results depend on the browsing harness and the date of the run, so compare both alongside scores.
View benchmarkDevTool Arena
Compares development environments and sandboxes through practical tasks. Results describe the tested infrastructure and workflow, so they are not standalone model capability scores.
View benchmarkDrivingBench
Puts a frontier model in control of a real Toyota Corolla's steering, throttle, and brakes through three MCP tools and scores progress along a fixed cone course. Compare attempt counts and harness latency before comparing progress.
View benchmarkGAIA
Tests assistants on questions requiring reasoning, information retrieval, and tool use. Difficulty levels and available tools influence the reported completion rate.
View benchmarkGDPval
OpenAI's evaluation of well-specified workplace tasks drawn from real occupations across nine economic sectors. Grading uses expert comparison against human deliverables, so read the win-rate definition before comparing models.
View benchmarkMLE-bench
Tests agents on machine-learning engineering tasks drawn from competitions. Results reflect the full workflow, including data preparation, training, and submission quality.
View benchmarkOSWorld
Tests agents on tasks in desktop applications. Evaluation checks the resulting computer state, with performance affected by the model’s interaction setup and task budget.
View benchmarkPostTrainBench
Tests whether CLI agents can post-train small base models under a fixed compute budget, scored on seven target benchmarks. Epoch AI verified the design, but compare the budget and harness alongside results.
View benchmarkSkillsBench
Tests how reusable skills affect agent task performance. Compare the skill configuration, model, and harness to distinguish improvements from changes in the setup.
View benchmarkToolBench
Evaluates models on solving tasks with external tools. Check the tool collection, evaluation method, and execution conditions before comparing reported results.
View benchmarkVending-Bench 2
Tests whether an agent can run a simulated vending machine business profitably and stay coherent over a full simulated year. Compare the run length and tool setup before comparing final balances.
View benchmarkWebArena
Tests agents on tasks across realistic websites. Success is graded against the required outcome, covering workflows that involve navigation and changing application state.
View benchmarkτ²-bench
Tests agents in tool-using conversations with simulated users. Success depends on following domain policies and reaching the required state across the interaction.
View benchmarkSecurity
Find, reproduce, and assess software vulnerabilities.
Cybench
Cybersecurity capture-the-flag challenges measuring autonomous vulnerability discovery and exploitation in sandboxed environments. Compare the guidance setting and agent scaffold before comparing solve rates.
View benchmarkCyberGym
Tests whether agents can reproduce real software vulnerabilities. Agents receive a description and unpatched codebase, then produce executable proof-of-concept inputs.
View benchmarkExploitBench
Measures how far agents can take real, hardened vulnerabilities toward a working exploit, scored on a ladder of verifiable capability tiers. Few exploits are public, which limits contamination but complicates reproduction.
View benchmarkExploitGym
Tests whether agents can turn known vulnerabilities into working exploits. Evaluation runs in controlled environments and distinguishes exploitation from simply reproducing a crash.
View benchmarkSEC-Bench Pro
Tests long-horizon security bug hunting in complex systems. Evaluation uses reproducible validation for tasks involving targets such as browser engines and the Linux kernel.
View benchmarkVision & design
Read images and charts, recognize objects, and judge visual output.
BabyVision
Tests visual discrimination, tracking, spatial perception, and pattern recognition. Direct vision answers and tool-assisted results use different setups and should be compared separately.
View benchmarkChartography
Tests understanding of real professional charts using expert-written questions. Charts span domains including science, finance, healthcare, engineering, and manufacturing.
View benchmarkDesignArena
Compares visual output through arena-style evaluations. Preference results depend on the task and judging setup, rather than a universal measure of design quality.
View benchmarkRoboflow Model Leaderboard
Compares computer-vision models on published evaluation tasks. Check the dataset, model configuration, and task-specific metric to interpret accuracy and performance results.
View benchmarkZeroBench
Tests difficult visual reasoning with separate main-question and subquestion scores. Compare sampling metrics and tool access before treating two reported results as equivalent.
View benchmarkEmbeddings & retrieval
Represent meaning, find relevant content, and rank results.
MTEB
Evaluates embedding models across multiple language tasks. Compare the benchmark version, language coverage, and task subset, especially when choosing models for retrieval.
View benchmarkReranker Simple Benchmark
Compares reranking models on ordering retrieved content. Check the dataset, candidate set, and ranking metric to judge relevance to your search workflow.
View benchmarkSpeech & audio
Recognize speech and evaluate spoken interactions.
UltraEval-Audio
Evaluates audio models across supported speech and audio tasks. Results depend on the dataset and task-specific metric, so compare matching evaluation configurations.
View benchmarkModel & infrastructure comparisons
Compare broad capability, human preference, speed, and hardware fit.
Arena
Compares models using human preferences in paired interactions. Ratings reflect the voting population, category, and collection period, rather than objective correctness alone.
View benchmarkArtificial Analysis
Publishes model evaluations and performance comparisons. Read each metric’s methodology and tested configuration when comparing capability, speed, cost, or human preference.
View benchmarkEpoch AI Benchmarking Hub
Epoch AI's hub of independently run benchmark results, the Epoch Capabilities Index, and benchmark reviews rating evaluations Verified or Flawed. Read each review before trusting a benchmark's headline numbers.
View benchmarkInferenceX
Compares inference performance across model and hardware configurations. Throughput and latency depend on the workload, serving stack, and resource allocation used.
View benchmarkLiveBench
Evaluates models on regularly refreshed tasks across several capabilities. Compare the release date and task category because the question set changes over time.
View benchmarkMETR Time Horizons
Estimates the length of software tasks a model completes correctly more often than not. Horizons are model estimates from a fixed task suite, so compare the suite version and confidence intervals.
View benchmarkOpen Weights
Helps compare openly available models. Check the evidence behind each comparison, including the model version, evaluation source, and deployment configuration.
View benchmarkWhichLLM
Helps explore model choices through published comparisons. Verify the source and date of each result before using it to select a model for your workload.
View benchmarkSpecialist tasks & behavior
Legal work, ethical judgment, uncertainty, games, and constrained training.
AI Battlegrounds
Evaluates agents through game-based competition. Outcomes depend on the game rules, opponents, and agent configuration rather than a general measure of model capability.
View benchmarkBullshitBench
Tests how models respond to flawed or nonsensical prompts. It examines whether a model challenges the premise instead of confidently inventing an answer.
View benchmarkHarvey LAB (Legal Agent Benchmark)
Evaluates agents on legal work. Read the task definitions and grading methodology to understand which parts of a legal workflow the results support.
View benchmarkKellyBench
Tests long-horizon decision-making in simulated sports betting markets. Agents build models, size bets, manage risk, and adapt their strategy across a football season.
View benchmarkParameter Golf
Compares language-model training approaches under tight resource constraints. Results reflect the training rules and model-size budget as well as the final evaluation score.
View benchmarkPhilosophy Bench
Evaluates model responses to philosophical questions. Interpretation depends on the prompts and judging methodology, rather than a single objective notion of philosophical correctness.
View benchmarkVals Minecraft Research Run
Reports progress from one 141-hour GPT-6 Astra Minecraft run in Normal Survival with keepInventory enabled. These observations are not a standardized cross-model leaderboard or speed ranking.
View benchmarkWhat to look for
- 01Does the benchmark test the task you care about? Multiple-choice reasoning tells you nothing about multi-file refactoring.
- 02Is contamination plausible? Public benchmarks leak into training data and scores drift upward without capability doing the same.
- 03How is success graded: tests, exact answers, human review, or a model judge? Read the grading rules and known limitations.
Common questions
- Where are the benchmarks from the DeepSeek V4.1 Flash launch?
- All 19 launch evaluations are represented here by 16 benchmark entries. Terminal-Bench 2.1, 3.0, and 4.0 share one entry; HLE with and without tools share Humanity's Last Exam. DeepSWE v1.1, GPQA Diamond, MathArena Apex, and ZeroBench-main retain their specific scope. The Model Atlas carries the reported scores, harness details, and source discrepancies.
- What is SWE-bench?
- A benchmark built from real GitHub issues and their merged fixes. The model gets the repository and the issue, and its patch is scored by whether the project's own tests pass.
- Why does the top benchmark model not feel best in practice?
- Benchmarks measure isolated task completion under ideal context. Daily work is about following your conventions, handling ambiguity, and staying useful over a long session — none of which is scored.
- Is there a benchmark for my specific domain?
- Increasingly, yes. Legal reasoning, ethical dilemmas, embeddings, object detection, speech and React code all have dedicated benchmarks now, and a narrow one measuring your actual task usually tells you more than a general leaderboard does.
More in Understand the AI landscape
- Evaluate an LLM applicationBuild test sets, score outputs, and catch quality regressions.
- Observe an LLM applicationTrace calls, inspect failures, and monitor latency, quality, and spend.
- Choose an inference providerCompare model routers, inference clouds, cloud catalogs, and direct lab APIs without collapsing them into one category.
- Run models locallyUse local inference runtimes and model managers on your own hardware.
- Add vector searchStore embeddings and retrieve relevant context for AI applications.