Capital & Compute
Directory· Updated July 27, 2026

AI benchmarks

Every model launch quotes a wall of benchmark names. This directory maps 116 of them across 12 categories, from the coding and agent tests builders watch most to reasoning, math, knowledge, instruction following, long context, multimodal, professional domains, human preference, safety and security: what each one measures, who built it, the year, how it is scored, a representative current top score, and a link to its leaderboard. 27 of them have a full write-up of their own. For models ranked by value, see the value leaderboard; for why these scores are easier to trust some years than others, our guide on whether AI benchmarks are reliable.

What are the main AI benchmarks?

The most-watched AI benchmarks in 2026, by what they test, are:

  • SWE-bench Verified and DeepSWE: real-world coding agents
  • Terminal-Bench and Frontier-Bench: agentic terminal and senior engineering tasks
  • OSWorld 2.0 and AutomationBench: computer-use and business-workflow agents
  • ARC-AGI-3: interactive reasoning, and the widest human-model gap left
  • Humanity’s Last Exam and GPQA Diamond: expert breadth and PhD-level science
  • FrontierMath and AIME: research and competition math
  • MMLU-Pro: broad academic knowledge
  • LMArena: human-preference ranking (Elo)
  • MMMU: college-level multimodal understanding

Older sets like MMLU, GSM8K and HumanEval are now saturated, with top models above 95%, so they are quoted mainly out of habit.

Which benchmarks evaluate AI agent reliability?

The benchmarks built to evaluate agent reliability in 2026 are tau-bench (whether a tool-using agent completes multi-turn tasks consistently, measured pass^k across repeated runs), Terminal-Bench (hard end-to-end command-line tasks), OSWorld 2.0 (computer-use agents on long-horizon desktop work), AutomationBench (cross-application business workflows graded on end state), GAIA and AgentBench (multi-step assistant tasks), and WebArena (long-horizon web tasks). For live rankings from real usage rather than a fixed test set, see Agent Arena; for what the scores mean in practice, our guide to AI agent benchmarks in 2026.

How many AI benchmarks are there?

There is no fixed number: hundreds of AI benchmarks exist, and new ones appear every month as older ones saturate. What is countable is the set frontier labs actually report. This directory maps 116 of them across 12 categories, of which28 are already saturated and 10 were released in 2026 alone. The practical figure to hold onto is smaller still: roughly a dozen benchmarks carry most of the signal in any given model launch, and they are listed above.

What is the hardest AI benchmark in 2026?

By the size of the remaining human-model gap, ARC-AGI-3: humans scored 100% at its March 2026 launch against 0.51% for frontier models, and the ARC-Prize-verified top is still only about 30%. Among knowledge benchmarks the hardest is Humanity’s Last Exam at roughly 53%. For long-horizon agent work, OSWorld 2.0 sits near 20% full task completion. Anything where frontier models score above 90% is measuring its own ceiling instead of the model.

AI benchmarks vs MLPerf: which one do you mean?

The phrase covers two different things. Model benchmarks, the subject of this directory, measure what a model can do: accuracy, task completion, reasoning, reliability. Systems benchmarks such as MLPerf measure how fast and how efficiently hardware runs a fixed workload, reported as throughput, latency and energy rather than capability. A model benchmark tells you which model to pick; a systems benchmark tells you what it costs to serve. For the cost side of that question, see cost per task and the value leaderboard.

116
Benchmarks mapped
Across 12 capability categories
25
Coding & agent tests
The deepest category, and the one builders watch most
28
Now saturated
Top models near the ceiling, so the score no longer separates them
10
Released in 2026
Harder replacements built after the previous generation saturated

What is an AI benchmark?

An AI benchmark is a standardized test, made of a fixed dataset, a task specification, and a scoring metric, used to measure and compare how well AI models perform a specific skill such as reasoning, coding, math, or knowledge. Running many models through the same test produces a single comparable number, which is what a model leaderboard ranks. The catch is that a benchmark only stays meaningful while it is hard: once the frontier clears it, or its answers leak into training data, the score stops telling good models from great ones, and the field has to build a harder one.

How long does an AI benchmark stay useful?

Shorter every year. A benchmark is useful only while the frontier has not cleared it, and that window is collapsing: ARC-AGI-1 survived six years from its 2019 release, the 2021 cohort (HumanEval, GSM8K, MMLU) held for roughly three, the 2023 to 2024 cohort for about two, and ARC-AGI-2 went from 54% to an ARC-Prize-verified 92.5% inside a single year. The trend is not perfectly monotonic, and GPQA Diamond outlasting SWE-bench by a year is a real exception rather than noise we have smoothed away. But the direction is unambiguous, it is the mechanism behind the 28 saturated rows in this directory, and it is why a benchmark name in a launch table means nothing without a date attached to it.

How long each AI benchmark stayed usefulA timeline chart of how long each benchmark stayed useful, from release year to saturation year: HumanEval released 2021, saturated 2024; GSM8K released 2021, saturated 2024; MMLU released 2021, saturated 2024; ARC-AGI-1 released 2019, saturated 2025; SWE-bench released 2023, saturated 2025; GPQA Diamond released 2023, saturated 2026; OSWorld released 2024, saturated 2026; SWE-bench Verified released 2024, saturated 2026; FrontierMath released 2024, saturated 2026; ARC-AGI-2 released 2025, saturated 2026; MMLU-Pro released 2024, not yet saturated; Humanity's Last Exam released 2025, not yet saturated; ARC-AGI-3 released 2026, not yet saturated.Useful life, then saturatedStill separates frontier models20192020202120222023202420252026HumanEval3 yrsGSM8K3 yrsMMLU3 yrsARC-AGI-16 yrsSWE-bench2 yrsGPQA Diamond3 yrsOSWorld2 yrsSWE-bench Verified2 yrsFrontierMath2 yrsARC-AGI-21 yrMMLU-Pro2 yrs+Humanity's Last Exam1 yr+ARC-AGI-3NewRelease year to saturation year (Capital & Compute assessment)
How long each AI benchmark stayed useful
BenchmarkReleasedSaturatedUseful life
HumanEval202120243 yrs
GSM8K202120243 yrs
MMLU202120243 yrs
ARC-AGI-1201920256 yrs
SWE-bench202320252 yrs
GPQA Diamond202320263 yrs
OSWorld202420262 yrs
SWE-bench Verified202420262 yrs
FrontierMath202420262 yrs
ARC-AGI-2202520261 yr
MMLU-Pro2024Not yet saturated2 yrs+
Humanity's Last Exam2025Not yet saturated1 yr+
ARC-AGI-32026Not yet saturatedNew
Useful life of thirteen benchmarks, from release to saturation. ARC-AGI-1 lasted six years, the 2021 cohort about three, and ARC-AGI-2 one. Green spans are still open, meaning the benchmark has not yet been cleared: ARC-AGI-3 and Humanity's Last Exam are where the headroom now is.Source: Capital & Compute AI benchmark directory, July 27, 2026. Saturation years are our assessment, set from each benchmark's own leaderboard trajectory, not a figure any benchmark publishes.

Benchmarks by category

The 116 benchmarks split into 12 categories. Coding and agents is the largest, both because it is the most commercially watched skill and because contamination forced builders to keep replacing the older sets. To search or filter all of them at once, use the directory tool below.

Underlined namesopen a full guide to that benchmark: how it works, its saturation history, and what its score hides. 27 of the 116 have one.

AI benchmarks tracked, by categoryA horizontal bar chart of how many benchmarks sit in each category: coding 25, agents & tool use 19, reasoning 11, mathematics 8, knowledge 7, instructions & languages 5, long context 8, multimodal 9, professional domains 5, human preference 12, truthfulness 3, security 4.0510152025Coding25Agents & tool use19Reasoning11Mathematics8Knowledge7Instructions & languages5Long context8Multimodal9Professional domains5Human preference12Truthfulness3Security4
AI benchmarks tracked, by category
ItemValue
Coding25
Agents & tool use19
Reasoning11
Mathematics8
Knowledge7
Instructions & languages5
Long context8
Multimodal9
Professional domains5
Human preference12
Truthfulness3
Security4
How the tracked benchmarks split across the nine categories. Coding and software-engineering agents is the deepest, reflecting both demand and the churn of contamination-driven replacements.Source: Capital & Compute AI benchmark directory, July 27, 2026

Coding & software-engineering agents

Whether a model can write, edit, and fix real code, increasingly as a multi-step agent working in a real repository. The most-watched category for AI builders, and where contamination bites hardest.

BenchmarkWhat it measuresStatusTop score
Aider PolyglotAider · 2024Full guideHow well a model writes and correctly edits code across many languages, including applying diffs in the right format and self-correcting after test failures.Leaderboard ↗ActiveNot confirmedPercent correct after the second attempt, plus percent using the correct edit format
BigCodeBenchBigCode project · 2024Whether models can write code that correctly invokes multiple function calls from diverse real libraries to satisfy complex, practical instructions.Leaderboard ↗ActiveNot confirmedpass@1 against rigorous per-task test suites
Commit0Zhao, Jiang, Lee et al. · 2024Whether an agent can write an entire Python library from scratch against an API specification and an interactive test suite.Leaderboard ↗ActiveNot confirmed% of unit tests passed
CursorBenchAnysphere · 2026Whether a coding agent can handle ambiguous, multi-file requests inside a real repository, judged on solution correctness, code quality, efficiency and interaction behaviour.Leaderboard ↗ActiveNot confirmedAgentic graders scoring correctness plus quality dimensions, since the requests are underspecified and admit several valid solutions
DeepSWEDatacurve · 2026Whether frontier coding agents can complete original, long-horizon engineering tasks written from scratch, with no upstream PR to memorize.Leaderboard ↗Active73% (v1.1)GPT-5.6 Solpass@1 (committed code graded in a clean environment) · read 2026-07
EvalPlusLiu, Xia, Wang et al. · 2023The same function-synthesis task as HumanEval and MBPP, rescored against far larger automatically generated test suites.Leaderboard ↗SaturatedNot confirmedpass@1 under the extended tests
Frontier-BenchThe Terminal-Bench and Harbor team (Marten, Shaw, Konwinski) with 100+ task contributors and reviewers · 2026Whether a coding agent can do senior-level engineering work: building features from realistic instructions, investigating bugs that need runtime inspection, and shipping code that matches an existing repository's conventions.Leaderboard ↗Active34.4%GPT-5.6 SolResolution rate (mean reward over repeated attempts), reported alongside cost and token use · read 2026-07
FrontierCodeCognition · 2026Whether a coding agent produces a mergeable, production-quality pull request, not just one that passes tests, judged on correctness, regression safety, scope, tests and style.Active13.4% (Diamond)Claude Opus 4.8Pass rate on blocker criteria plus a weighted six-dimension quality rubric · read 2026-06
HumanEvalOpenAI · 2021Full guideWhether a model can synthesize a single correct Python function from a docstring so that it passes the provided unit tests.Saturated~99%Frontier models broadlypass@k (primarily pass@1) · read 2025-04
KernelBenchOuyang, Guo, Arora et al. · 2025Whether a model can write GPU kernels that are both correct and actually faster than the PyTorch baseline.ActiveNot confirmedfast_p: % of generated kernels that are correct and at least p times faster than baseline
LiveCodeBenchUC Berkeley, MIT and Cornell · 2024Full guideCode generation and related skills (self-repair, execution, test-output prediction) on fresh competitive-programming problems, designed to be contamination-free.Leaderboard ↗ActiveNot confirmedpass@1
MBPPGoogle Research · 2021Whether a model can generate short, entry-level Python functions from a natural-language prompt that pass the provided tests.Saturated~95%+Frontier models broadlypass@1 · read 2026-06
Multi-SWE-benchByteDance · 2025Cross-language issue resolution: whether agents can resolve real GitHub issues with a passing patch across many languages beyond Python.Leaderboard ↗ActiveNot confirmed% resolved (pass@1)
PaperBenchOpenAI · 2025Whether an agent can replicate a published AI research paper from scratch: understand the contribution, build the codebase, and run the experiments.ActiveNot confirmedReplication score against a hierarchical rubric, graded by an LLM judge
RepoBenchLiu, Xu and McAuley · 2023Repository-level code auto-completion: retrieving relevant cross-file context, predicting the next line, and the combined retrieval-plus-completion pipeline.ActiveNot confirmedRetrieval accuracy and exact-match / edit similarity for next-line completion
SciCodeTian, Gao, Zhang et al. · 2024Whether a model can write code that solves real scientific research problems, not general software tasks.Leaderboard ↗ActiveNot confirmed% of subproblems and main problems solved
SWE-benchPrinceton and Stanford · 2023Full guideWhether a system can resolve a real GitHub issue by generating a patch that passes the repository's hidden tests.Leaderboard ↗SaturatedNot confirmed% resolved (pass@1)
SWE-bench MultimodalStanford and Princeton · 2024Whether coding agents can resolve real GitHub issues in visual, user-facing JavaScript software where the bug or feature involves the UI.Leaderboard ↗ActiveNot confirmed% resolved (pass@1)
SWE-bench ProScale AI · 2025Full guideWhether agents can solve long-horizon, enterprise-grade software-engineering tasks under standardized scaffolding, designed to resist contamination.Leaderboard ↗Active59.1% (public set)GPT-5.4 (xHigh)% resolved (pass@1) under standardized agent scaffolding · read 2026-06
SWE-bench VerifiedOpenAI · 2024Full guideThe same real-GitHub-issue resolution task as SWE-bench, restricted to a human-validated subset where the issue is solvable and the tests are not broken.Leaderboard ↗Saturated~95%Claude Fable 5% resolved (pass@1) · read 2026-06
SWE-bench-LiveZhang, He, Zhang et al. · 2025The same real-GitHub-issue resolution task as SWE-bench, but on tasks harvested continuously from issues created after a model was trained.Leaderboard ↗ActiveNot confirmed% resolved (pass@1)
SWE-LancerOpenAI · 2025Whether frontier models can complete real paid freelance software jobs, both coding and technical-management tasks, well enough to earn the payouts.ActiveNot confirmedDollars earned (and % of tasks resolved)
SWE-PolyBenchRashid, Bock, Zhuang et al. · 2025Whether a coding agent can resolve repository-level tasks outside Python, across Java, JavaScript, TypeScript and Python.ActiveNot confirmed% resolved, plus syntax-tree-based retrieval and file-localisation metrics
SWE-rebenchBadertdinov, Golubev, Nekrashevich et al. · 2025Issue resolution on a continuously refreshed, decontaminated pool of Python software-engineering tasks mined automatically from open-source repositories.Leaderboard ↗ActiveNot confirmed% resolved (pass@1)
Terminal-BenchStanford and the Laude Institute · 2026Full guideWhether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal.Leaderboard ↗Active83.4% (v2.1)Codex (GPT-5.5)Pass/fail, graded by verification scripts in the agent's Docker environment (pass@1) · read 2026-06

Agents, tool use & computer use

Whether a model can plan, call tools, browse, and operate a computer or website to finish open-ended tasks, not just answer a question in one shot.

BenchmarkWhat it measuresStatusTop score
AgentBenchTsinghua University · 2023How well an LLM acts as an autonomous agent in multi-turn, open-ended decision-making across diverse interactive environments.Leaderboard ↗ActiveNot confirmedPer-environment success aggregated into an overall score
AndroidWorldRawles, Clinckemaillie, Chang et al. · 2024Whether an agent can operate a real Android phone to finish tasks across everyday apps.Leaderboard ↗ActiveNot confirmedProgrammatic reward from the device end state
AutomationBenchZapier · 2026Whether an agent can run a realistic business workflow end to end across several apps: discovering the right API endpoints itself, following a policy document, and writing correct data into every system it touches.Leaderboard ↗Active26.2%Claude Opus 5 (max effort)task_completed_correctly: strict pass/fail where every scored end-state assertion must pass, with partial credit reported only as a diagnostic · read 2026-07
BFCLUC Berkeley Gorilla team · 2024Full guideWhether a model calls functions and APIs correctly: picking the right function, filling parameters with valid types, and refusing to invent functions that were not offered.Leaderboard ↗ActiveNot confirmedAbstract-syntax-tree match against a reference call, plus executable checks; overall score is the unweighted mean of subcategories
BrowseCompOpenAI · 2025Whether a browsing agent can persistently navigate the open web to locate a single hard-to-find, entangled fact.Active51.5%OpenAI Deep Research (launch paper)Accuracy via model-graded semantic equivalence to the reference answer · read 2025-04
DeepSearchQAGoogle DeepMind · 2026Whether a deep-research agent can plan and execute a long chain of web searches to return an exhaustive, de-duplicated answer list rather than a single fact.Leaderboard ↗ActiveNot confirmedAccuracy against each task's objectively verifiable exhaustive answer set
GAIAMeta AI and Hugging Face · 2023Full guideWhether an AI assistant can answer real-world questions that require multi-step reasoning, multiple modalities, web browsing and general tool use.Leaderboard ↗Active~75%HAL agent (Claude Sonnet 4.5)Exact-match accuracy against an unambiguous answer · read 2026-06
LoCoMoMaharana, Lee, Tulyakov et al. · 2024Whether an agent remembers and reasons over a conversation that spans months, rather than a single session.ActiveNot confirmedQuestion-answering accuracy, event summarisation and multi-modal dialogue generation
Mind2Web 2Gou, Huang, Ning et al. · 2025Whether an agentic search or deep-research system can browse the live web and return a correct, citation-backed answer to a long-horizon question.Leaderboard ↗ActiveNot confirmedAgent-as-a-Judge rubric scoring of answer correctness and citation support
MLE-benchOpenAI · 2024Whether an AI agent can do end-to-end machine-learning engineering (data prep, training, experimentation, submission) at the level of human Kaggle competitors.Leaderboard ↗Active16.9% (paper baseline)o1-preview with AIDE scaffoldingMedal rate (fraction of competitions reaching bronze/silver/gold thresholds) · read 2024-10
OSWorldXLANG Lab, University of Hong Kong · 2024Whether a multimodal agent can operate a real computer (desktop apps, file I/O, multi-app workflows) to complete open-ended tasks in a live virtual machine.Leaderboard ↗SaturatedNot confirmedExecution-based success rate via per-task verification scripts that inspect machine state
OSWorld 2.0XLANG Lab, University of Hong Kong, with collaborators at Columbia, UCSB, UCSD, Mila, Ohio State and others · 2026Full guideWhether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state.Leaderboard ↗Active20.6%Claude Opus 4.8 (max thinking)Binary completion at a 500-step cap, reported with a weighted-checkpoint partial score · read 2026-06
tau-benchSierra · 2024Full guideWhether a tool-using agent can reliably complete customer-service tasks over multi-turn conversations with a simulated user while obeying domain policies.ActiveNot confirmedpass^k: the probability an agent succeeds across all k independent trials (reliability, not just average success)
TheAgentCompanyXu, Song, Li et al. · 2024Whether an agent can do real knowledge work inside a simulated software company: browsing, coding, using internal tools, and messaging simulated colleagues.Leaderboard ↗ActiveNot confirmedFull and partial task completion, scored by checkpoint
Vending-BenchAndon Labs · 2025Whether an agent stays coherent over a very long horizon, by running a simulated vending-machine business: stock, orders, pricing and daily fees.ActiveNot confirmedNet worth and units sold at the end of the run
VisualWebArenaCarnegie Mellon University · 2024Whether a multimodal agent can complete visually grounded web tasks that require interpreting images and page layout, not just text.Leaderboard ↗ActiveNot confirmedFunctional success rate via execution-based evaluation
WebArenaCarnegie Mellon University · 2023Whether an autonomous agent can complete long-horizon, realistic web tasks (navigation, forms, multi-step workflows) in fully functional self-hosted websites.Leaderboard ↗ActiveNot confirmedFunctional success rate via execution-based reward checking the end state
WebVoyagerHe, Yao, Ma et al. · 2024Whether a multimodal web agent can complete a user instruction end to end on real, live websites rather than a simulator or a static snapshot.SaturatedNot confirmedTask success rate, judged automatically from screenshots and responses
Windows Agent ArenaBonatti, Zhao, Bonacci et al. · 2024Whether a multimodal agent can operate a full Windows desktop across the applications people actually use at work.ActiveNot confirmedTask success rate from the OS end state

Reasoning & abstraction

Hard multi-step reasoning and fluid, novel problem-solving designed to resist memorization. The benchmarks the frontier is still far from solving.

BenchmarkWhat it measuresStatusTop score
AGIEvalZhong, Cui, Guo et al. · 2023Human-centric reasoning, using questions from real standardised exams taken by people rather than synthetic datasets.SaturatedNot confirmedAccuracy
ARC-AGI-1Francois Chollet · 2019Whether a system can infer the abstract rule of a novel visual grid puzzle from a few examples and apply it to a new input.Leaderboard ↗Saturated97.5% (public eval)Claude Opus 5 and GPT-5.6 Solpass@2 exact-grid-match accuracy · read 2026-07
ARC-AGI-2ARC Prize Foundation · 2025Full guideThe same fluid-intelligence test as ARC-AGI-1, but with harder, contamination-resistant tasks that stay easy for humans yet very hard for AI.Leaderboard ↗Saturated92.5% (semi-private)GPT-5.6 Sol (max effort)pass@2 exact-grid-match accuracy, reported with a cost-per-task efficiency metric · read 2026-07
ARC-AGI-3ARC Prize Foundation · 2026Full guideWhether an agent dropped into an unfamiliar interactive environment with no instructions, stated goal or rules can work out what to do by acting, build a usable world model, and keep learning across levels.Leaderboard ↗Active30.2% (public demo)Claude Opus 5 (high effort)Games beaten at or above human-level action efficiency, measuring skill-acquisition efficiency rather than one-shot accuracy · read 2026-07
BIG-Bench HardSuzgun et al. · 2022A suite of multi-step reasoning tasks (logic, arithmetic, algorithmic, commonsense) on which pre-2022 models trailed average human raters.SaturatedNot confirmedPer-task accuracy averaged across the 23 tasks
DROPDua, Wang, Dasigi et al. · 2019Reading comprehension that requires discrete operations over a passage: resolving references then adding, counting or sorting.SaturatedNot confirmedF1 and exact match
EnigmaEvalScale AI · 2025Long multimodal puzzle solving: finding hidden connections between unrelated pieces of information and chaining many deductive steps.Leaderboard ↗ActiveNot confirmedExact-match accuracy on the final puzzle answer
GPQA DiamondRein et al. · 2023Full guideGraduate and PhD-level multiple-choice scientific reasoning in biology, physics and chemistry, on questions designed to be unanswerable by quick web search.Leaderboard ↗Saturated~94%Gemini 3.1 Pro PreviewMultiple-choice accuracy (random baseline 25%, PhD-expert baseline about 70%) · read 2026-02
Humanity's Last ExamCenter for AI Safety (CAIS) and Scale AI · 2025Full guideFrontier, closed-ended expert knowledge and reasoning across more than 100 academic disciplines at the limit of human expertise.Leaderboard ↗Active53.3%Claude Fable 5 (Max Effort)Accuracy (exact match / multiple-choice), often reported with a calibration metric · read 2026-06
MuSRSprague, Ye, Durrett et al. · 2023Multistep commonsense reasoning embedded in long natural-language narratives such as murder mysteries, object placement and team allocation.Leaderboard ↗ActiveNot confirmedMultiple-choice accuracy
ZebraLogicLin, Le Bras, Richardson et al. · 2025Logical deduction under hard constraints, using logic grid puzzles generated from constraint-satisfaction problems.ActiveNot confirmedPuzzle-level accuracy (all cells correct)

Mathematics

From grade-school word problems to research-level proofs. The older sets are saturated; the newest are held back from the public to stay contamination-resistant.

BenchmarkWhat it measuresStatusTop score
AIME 2025Mathematical Association of America; adopted as an LLM eval by the community · 2025Olympiad-track competition mathematics at the level of the American Invitational Mathematics Examination, used as a high-difficulty LLM eval.Leaderboard ↗Saturated100%Multiple frontier reasoning modelsExact-match accuracy, usually pass@1 averaged over samples · read 2026-06
FrontierMathEpoch AI · 2024Full guideResearch-level original mathematics requiring hours to days of expert effort, across number theory, analysis, algebraic geometry and more.Leaderboard ↗Saturated87% (Tiers 1-3)Claude Fable 5Accuracy (fraction with a correct, automatically verifiable final answer) · read 2026-06
GSM8KOpenAI · 2021Multi-step grade-school arithmetic word-problem reasoning.Leaderboard ↗Saturated~99.6%Frontier models broadlyExact-match accuracy on the final numeric answer · read 2026-05
MATHHendrycks et al. · 2021Step-by-step solving of high-school competition mathematics across algebra, geometry, number theory, probability and precalculus.Leaderboard ↗Saturated~99% (MATH-500)GPT-5Exact-match accuracy on the final boxed answer · read 2026-04
MathArenaETH Zurich · 2025Mathematical reasoning and proof-writing on freshly released competition problems, evaluated before they can enter training data.Leaderboard ↗Active81.1% (aggregate)GPT-5.5 (xhigh)Per-competition accuracy and an aggregate expected-performance score · read 2026-04
miniF2FZheng, Han and Polu · 2021Formal theorem proving on Olympiad-level mathematics, as a shared benchmark across proof assistants.SaturatedNot confirmed% of statements formally proved
Omni-MATHGao, Song, Cai et al. · 2024Olympiad-level mathematical reasoning across a broad range of subdomains and difficulty levels.Leaderboard ↗ActiveNot confirmedAccuracy, scored with an LLM-based verifier (Omni-Judge)
PutnamBenchTsoukalas, Lee, Jennings et al. · 2024Whether a neural theorem prover can produce a formal, machine-checked proof of an undergraduate competition problem.Leaderboard ↗ActiveNot confirmed% of theorems formally proved and machine-verified

Knowledge & general QA

Broad academic and factual knowledge across domains, usually multiple-choice. The most-quoted and most-saturated family, now largely replaced by harder variants.

BenchmarkWhat it measuresStatusTop score
HellaSwagZellers, Holtzman, Bisk et al. · 2019Commonsense sentence completion: picking the plausible continuation of an everyday scenario.RetiredNot confirmedAccuracy
MMLUHendrycks et al. · 2021Full guideBroad academic and professional knowledge across 57 subjects via four-choice multiple-choice questions.Leaderboard ↗Saturated~93%Qwen3.7 MaxAccuracy · read 2026-06
MMLU-ProTIGER-Lab · 2024Full guideHarder multi-task reasoning and knowledge designed to de-saturate MMLU and reward deliberate reasoning over recall.Leaderboard ↗Active~90%Gemini 3 Pro PreviewAccuracy · read 2026-06
MMLU-ReduxGema et al. · 2024A re-annotated, error-corrected subset of MMLU used to measure true knowledge accuracy without the original's label noise.Leaderboard ↗ActiveNot confirmedAccuracy on cleaned labels
SimpleQAOpenAI · 2024Short-form parametric factuality: whether a model answers single-answer fact-seeking questions correctly and abstains when unsure.Leaderboard ↗ActiveNot confirmedAccuracy, plus correct-given-attempted and an F-score balancing attempts against accuracy
SuperGPQAM-A-P Team · 2025Graduate-level knowledge and reasoning across 285 disciplines, including the applied and service fields that mainstream benchmarks ignore.ActiveNot confirmedAccuracy
TriviaQAJoshi, Choi, Weld et al. · 2017Factual recall and reading comprehension over trivia questions with evidence documents.RetiredNot confirmedExact match and F1

Instruction following & multilingual

Whether a model does exactly what it was told, and whether it does it as well outside English. Two skills that decide production reliability but rarely make a launch slide.

BenchmarkWhat it measuresStatusTop score
Global-MMLUSingh, Romanou, Fourrier et al. · 2024Multilingual academic knowledge, separating questions that are culturally neutral from those requiring culture-specific knowledge.ActiveNot confirmedAccuracy
IFEvalZhou, Lu, Mishra et al. · 2023Whether a model obeys instructions that can be checked by a program, such as a word count, a required keyword, or a forbidden format.SaturatedNot confirmedStrict and loose instruction-following accuracy
INCLUDERomanou, Foroutan, Sotnikova et al. · 2024Multilingual understanding built from local exam material, so the questions test regional knowledge rather than translated Western content.ActiveNot confirmedAccuracy
MGSMShi, Suzgun, Freitag et al. · 2022Grade-school math word problems solved via chain-of-thought reasoning in ten languages.SaturatedNot confirmedAccuracy
Multi-IFHe, Jin, Wang et al. · 2024Whether a model keeps following instructions across multiple turns and in languages other than English.ActiveNot confirmedInstruction-following accuracy per turn

Long context & retrieval

Whether a model can actually use a very long input, not just accept it: finding facts, resolving references, and reasoning across hundreds of thousands of tokens.

BenchmarkWhat it measuresStatusTop score
BABILongKuratov, Bulatov, Anokhin et al. · 2024Reasoning over facts deliberately scattered through an extremely long document, not just retrieving one of them.ActiveNot confirmedAccuracy by context length
HELMETYen, Gao, Hou et al. · 2024Long-context ability across a wide spread of realistic downstream applications rather than one synthetic retrieval task.Leaderboard ↗ActiveNot confirmedPer-category task metrics, reported across context lengths
LOFTLee, Chen, Dai et al. · 2024Whether a long-context model can replace a retrieval pipeline outright: doing retrieval, RAG and SQL-style tasks natively from context.ActiveNot confirmedTask-specific accuracy compared against specialised retrieval pipelines
LongBenchTsinghua University · 2023Comprehensive long-context understanding across realistic tasks (QA, summarization, few-shot, code, synthetic) in English and Chinese.Leaderboard ↗Active57.7% (v2, with reasoning)o1-previewv1: per-task automatic metrics. v2: multiple-choice accuracy · read 2024-12
MRCRGoogle DeepMind (Michelangelo); open-source variant by OpenAI · 2024Whether a model can distinguish and retrieve the correct one among multiple near-identical requests buried in a long multi-turn conversation.ActiveNot confirmedSimilarity of the model’s output to the target instance, gated by a required answer-prefix
Needle-in-a-HaystackGreg Kamradt · 2023Whether a model can recall a single planted fact (the needle) inserted at varying depths within a long context (the haystack).SaturatedNot confirmedRetrieval accuracy at each depth and length cell
NoLiMaAdobe Research and LMU Munich · 2025Long-context retrieval and reasoning when the question and the target fact share minimal literal word overlap, forcing latent association rather than keyword matching.Leaderboard ↗ActiveNot confirmedAccuracy at each length, relative to the model's short-context baseline
RULERNVIDIA · 2024Full guideThe real effective context length of a model by testing retrieval, multi-hop tracing, aggregation and QA at increasing sequence lengths.Leaderboard ↗ActiveNot confirmedWeighted-average accuracy across tasks and lengths; effective length is the longest length still above threshold

Multimodal & vision

Reasoning over images, charts, documents, and video alongside text. The frontier for models that see, not just read.

BenchmarkWhat it measuresStatusTop score
ChartQAMasry, Long, Tan et al. · 2022Question answering over charts that requires both reading visual features and doing arithmetic on them.SaturatedNot confirmedRelaxed accuracy (numeric answers within a tolerance)
CharXivWang, Xia, He et al. · 2024Chart understanding on real, messy scientific figures rather than clean template-generated charts.Leaderboard ↗ActiveNot confirmedAccuracy, split into descriptive and reasoning questions
DocVQAMathew, Karatzas and Jawahar · 2020Question answering over scanned document images, where layout and structure carry the meaning.SaturatedNot confirmedANLS (average normalised Levenshtein similarity)
MathVistaLu et al. · 2023Mathematical and quantitative reasoning grounded in visual contexts such as figures, charts, geometry and scientific diagrams.Leaderboard ↗Active~91% (testmini)Seed 2.1 ProAccuracy · read 2026-06
MMBenchLiu, Duan, Zhang et al. · 2023Fine-grained vision-language ability across a structured taxonomy of perception and reasoning skills.Leaderboard ↗SaturatedNot confirmedAccuracy under circular evaluation
MMMUMMMU team · 2023Full guideCollege-level multimodal understanding and reasoning over images, diagrams, charts and text across many disciplines.Leaderboard ↗Active~86%Qwen3.6 PlusAccuracy · read 2026-06
MMMU-ProMMMU team · 2024A harder, contamination-resistant version of MMMU that forces genuine visual reasoning rather than text-only shortcuts.Leaderboard ↗Active~84%Gemini 3.5 FlashAccuracy · read 2026-06
MMStarChen, Li, Dong et al. · 2024Genuinely vision-dependent multimodal ability, on samples selected so the answer cannot be inferred from the text alone.ActiveNot confirmedAccuracy, reported alongside a multimodal-gain and multimodal-leakage measure
Video-MMEMME-Benchmarks team · 2024Comprehensive video understanding by multimodal LLMs across short, medium and long clips.Leaderboard ↗Active~89%Seed 2.1 ProAccuracy (tested with and without subtitles) · read 2026-06

Domain & professional work

Medicine, law, finance and other expert fields, where the failure cost is real and a general-purpose score tells you almost nothing about whether the model is safe to deploy.

BenchmarkWhat it measuresStatusTop score
FinanceBenchPatronus AI · 2023Open-book financial question answering over real public-company filings, with the supporting evidence required.ActiveNot confirmedAnswer correctness against the evidence, human reviewed
HealthBenchOpenAI · 2025Full guideOpen-ended clinical conversation quality and safety, graded against rubrics written by practising physicians.ActiveNot confirmedRubric score, graded by a model grader against physician-written criteria
LegalBenchGuha, Nyarko, Ho et al. · 2023Full guideLegal reasoning across the specific skills lawyers actually use, as defined by legal professionals.Leaderboard ↗ActiveNot confirmedPer-task accuracy, aggregated by reasoning type
MedHELMStanford CRFM · 2025Clinical ability across the breadth of real medical work, on a clinician-validated taxonomy rather than exam questions.Leaderboard ↗ActiveNot confirmedPer-task clinical metrics plus head-to-head win rates
MedQAJin, Pan, Oufattole et al. · 2020Medical knowledge, using real questions from professional medical board examinations.SaturatedNot confirmedAccuracy

Human preference & holistic

Aggregate and head-to-head measures: human-voted arenas, composite indices, and multi-metric frameworks that rank overall capability rather than one skill.

BenchmarkWhat it measuresStatusTop score
AlpacaEval 2 (Length-Controlled)Dubois, Galambosi, Liang et al. · 2024Instruction-following quality judged by an LLM, with a regression correction for the judge’s bias toward longer answers.Leaderboard ↗SaturatedNot confirmedLength-controlled win rate
Arena-Hard-AutoLi, Chiang, Frick et al. · 2024Human-preference-aligned quality on hard open-ended prompts, scored automatically instead of by live human voting.ActiveNot confirmedWin rate against a baseline model, judged by an LLM
Artificial Analysis Coding Agent IndexArtificial Analysis · 2026Overall coding-agent capability as a single number, scoring the full stack (a specific model plus its harness and settings) rather than a model in isolation.Leaderboard ↗ActiveNot confirmedSimple average of the component benchmark scores, with every task equally weighted
Artificial Analysis Intelligence IndexArtificial Analysis · 2024A composite index of overall model intelligence aggregating performance across reasoning, coding, knowledge, science and agentic tasks.Leaderboard ↗Active~60 (index)Claude Fable 5Composite index score (0 to 100 aggregate) · read 2026-06
Copilot ArenaChi, Chen, Angelopoulos et al. · 2025Which coding model developers actually prefer, collected from paired completions inside a real editor rather than a chat window.ActiveNot confirmedElo-style ranking from in-editor pairwise preferences
Epoch Capabilities IndexEpoch AI · 2025Full guideOverall model capability on one continuous scale, stitched together from many benchmarks of differing difficulty.Leaderboard ↗ActiveNot confirmedECI score on the anchored scale
GDPvalOpenAI, with an agentic re-run by Artificial Analysis as GDPval-AA · 2025Whether a model can produce the actual deliverables of skilled professional work (documents, slides, spreadsheets, diagrams) well enough to stand against an industry expert's version.Leaderboard ↗Active1861 Elo (GDPval-AA v2)Claude Opus 5 (adaptive reasoning, max effort)Blind pairwise comparison of two anonymised outputs on the same task, aggregated into an Elo rating; the v2 scale anchors human expert deliverables at 1000 · read 2026-07
HELMStanford CRFM · 2022Multi-metric holistic evaluation across many scenarios, reporting accuracy alongside calibration, robustness, fairness, bias, toxicity and efficiency.Leaderboard ↗ActiveNot confirmedMulti-metric (per-metric scores across scenarios; no single headline number)
LiveBenchWhite, Dooley, Roberts et al. · 2024Full guideBroad capability across six categories at once (math, coding, reasoning, data analysis, instruction following, language), on questions refreshed monthly.Leaderboard ↗ActiveNot confirmedObjective automatic scoring against ground truth, averaged across categories
LMArenaArena · 2023Full guideCrowdsourced human preference between two anonymized model responses, aggregated into a relative ranking, not an objective capability.Leaderboard ↗Active~1510 EloClaude Opus 4.8Elo / Bradley-Terry pairwise rating (an Arena Score) · read 2026-06
METR Time HorizonMETR · 2025Full guideModel capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success.ActiveNot confirmed50%-task-completion time horizon, in minutes or hours
MT-BenchLMSYS · 2023Instruction-following and conversational quality on multi-turn prompts, scored automatically by a strong LLM judge.Leaderboard ↗SaturatedNot confirmedLLM-as-judge score (1 to 10 scale, averaged)

Safety, hallucination & factuality

Whether a model tells the truth and resists making things up. Measures honesty and hallucination rate, not raw capability.

BenchmarkWhat it measuresStatusTop score
HaluEvalLi et al. · 2023A model's ability to recognize hallucinated content across question answering, knowledge-grounded dialogue and summarization.ActiveNot confirmedHallucination-recognition accuracy (faithful vs hallucinated)
TruthfulQALin, Hilton, Evans · 2021Full guideWhether a model avoids repeating common human misconceptions when answering questions, rather than imitating popular falsehoods.ActiveNot confirmed% truthful (and % truthful-and-informative)
Vectara Hallucination LeaderboardVectara · 2023How often a model introduces unsupported content when summarizing a provided source document, i.e. faithfulness in closed-book summarization.Leaderboard ↗Active1.8% (lower is better)antgroup/finix-s1-32bHallucination rate (% of summaries judged unfaithful; lower is better) · read 2026-05

Security & dangerous capability

Offensive cyber skill, hazardous knowledge, and whether an agent can be talked into doing harm. Here a high score is bad news, which inverts how every other category on this page reads.

BenchmarkWhat it measuresStatusTop score
AgentHarmAndriushchenko, Souly, Dziemian et al. · 2024Whether a tool-using agent refuses explicitly malicious multi-step tasks, and whether it stays capable enough to complete them once jailbroken.ActiveNot confirmedRefusal rate and post-jailbreak task-completion rate
CybenchZhang, Perry, Dulepet et al. · 2024Whether an agent can autonomously solve professional capture-the-flag security tasks: finding a vulnerability and executing an exploit.Leaderboard ↗ActiveNot confirmed% of tasks and subtasks solved unassisted
CyberSecEval 3Meta · 2024Cybersecurity risk across eight areas, split between risk to third parties and risk to the developers and users of an application.ActiveNot confirmedPer-risk pass and failure rates, measured with and without guardrails
WMDPLi, Pan, Gopal et al. · 2024Proxy knowledge of hazardous biosecurity, cybersecurity and chemical-security material.Leaderboard ↗ActiveNot confirmedAccuracy, where lower is the desired direction after unlearning

Search and filter all 116 benchmarks

Filter by category or status, or search by name, alias, what a benchmark measures, or who built it.

Search the directory

Filter all 116 benchmarks by category or status, or search by name, alias, what it measures, or who built it.

116 shown

Underlined names open a full guide to that benchmark. 27 of 116 have one.

BenchmarkWhat it measuresStatusTop score
Aider Polyglot Aider · 2024 Coding· Full guideHow well a model writes and correctly edits code across many languages, including applying diffs in the right format and self-correcting after test failures. Leaderboard ↗ActiveNot confirmed Percent correct after the second attempt, plus percent using the correct edit format
BigCodeBench BigCode project · 2024 CodingWhether models can write code that correctly invokes multiple function calls from diverse real libraries to satisfy complex, practical instructions. Leaderboard ↗ActiveNot confirmed pass@1 against rigorous per-task test suites
Commit0 Zhao, Jiang, Lee et al. · 2024 CodingWhether an agent can write an entire Python library from scratch against an API specification and an interactive test suite. Leaderboard ↗ActiveNot confirmed % of unit tests passed
CursorBench Anysphere · 2026 CodingWhether a coding agent can handle ambiguous, multi-file requests inside a real repository, judged on solution correctness, code quality, efficiency and interaction behaviour. Leaderboard ↗ActiveNot confirmed Agentic graders scoring correctness plus quality dimensions, since the requests are underspecified and admit several valid solutions
DeepSWE Datacurve · 2026 CodingWhether frontier coding agents can complete original, long-horizon engineering tasks written from scratch, with no upstream PR to memorize. Leaderboard ↗Active73% (v1.1) GPT-5.6 Sol pass@1 (committed code graded in a clean environment) · read 2026-07
EvalPlus Liu, Xia, Wang et al. · 2023 CodingThe same function-synthesis task as HumanEval and MBPP, rescored against far larger automatically generated test suites. Leaderboard ↗SaturatedNot confirmed pass@1 under the extended tests
Frontier-Bench The Terminal-Bench and Harbor team (Marten, Shaw, Konwinski) with 100+ task contributors and reviewers · 2026 CodingWhether a coding agent can do senior-level engineering work: building features from realistic instructions, investigating bugs that need runtime inspection, and shipping code that matches an existing repository's conventions. Leaderboard ↗Active34.4% GPT-5.6 Sol Resolution rate (mean reward over repeated attempts), reported alongside cost and token use · read 2026-07
FrontierCode Cognition · 2026 CodingWhether a coding agent produces a mergeable, production-quality pull request, not just one that passes tests, judged on correctness, regression safety, scope, tests and style. Active13.4% (Diamond) Claude Opus 4.8 Pass rate on blocker criteria plus a weighted six-dimension quality rubric · read 2026-06
HumanEval OpenAI · 2021 Coding· Full guideWhether a model can synthesize a single correct Python function from a docstring so that it passes the provided unit tests. Saturated~99% Frontier models broadly pass@k (primarily pass@1) · read 2025-04
KernelBench Ouyang, Guo, Arora et al. · 2025 CodingWhether a model can write GPU kernels that are both correct and actually faster than the PyTorch baseline. ActiveNot confirmed fast_p: % of generated kernels that are correct and at least p times faster than baseline
LiveCodeBench UC Berkeley, MIT and Cornell · 2024 Coding· Full guideCode generation and related skills (self-repair, execution, test-output prediction) on fresh competitive-programming problems, designed to be contamination-free. Leaderboard ↗ActiveNot confirmed pass@1
MBPP Google Research · 2021 CodingWhether a model can generate short, entry-level Python functions from a natural-language prompt that pass the provided tests. Saturated~95%+ Frontier models broadly pass@1 · read 2026-06
Multi-SWE-bench ByteDance · 2025 CodingCross-language issue resolution: whether agents can resolve real GitHub issues with a passing patch across many languages beyond Python. Leaderboard ↗ActiveNot confirmed % resolved (pass@1)
PaperBench OpenAI · 2025 CodingWhether an agent can replicate a published AI research paper from scratch: understand the contribution, build the codebase, and run the experiments. ActiveNot confirmed Replication score against a hierarchical rubric, graded by an LLM judge
RepoBench Liu, Xu and McAuley · 2023 CodingRepository-level code auto-completion: retrieving relevant cross-file context, predicting the next line, and the combined retrieval-plus-completion pipeline. ActiveNot confirmed Retrieval accuracy and exact-match / edit similarity for next-line completion
SciCode Tian, Gao, Zhang et al. · 2024 CodingWhether a model can write code that solves real scientific research problems, not general software tasks. Leaderboard ↗ActiveNot confirmed % of subproblems and main problems solved
SWE-bench Princeton and Stanford · 2023 Coding· Full guideWhether a system can resolve a real GitHub issue by generating a patch that passes the repository's hidden tests. Leaderboard ↗SaturatedNot confirmed % resolved (pass@1)
SWE-bench Multimodal Stanford and Princeton · 2024 CodingWhether coding agents can resolve real GitHub issues in visual, user-facing JavaScript software where the bug or feature involves the UI. Leaderboard ↗ActiveNot confirmed % resolved (pass@1)
SWE-bench Pro Scale AI · 2025 Coding· Full guideWhether agents can solve long-horizon, enterprise-grade software-engineering tasks under standardized scaffolding, designed to resist contamination. Leaderboard ↗Active59.1% (public set) GPT-5.4 (xHigh) % resolved (pass@1) under standardized agent scaffolding · read 2026-06
SWE-bench Verified OpenAI · 2024 Coding· Full guideThe same real-GitHub-issue resolution task as SWE-bench, restricted to a human-validated subset where the issue is solvable and the tests are not broken. Leaderboard ↗Saturated~95% Claude Fable 5 % resolved (pass@1) · read 2026-06
SWE-bench-Live Zhang, He, Zhang et al. · 2025 CodingThe same real-GitHub-issue resolution task as SWE-bench, but on tasks harvested continuously from issues created after a model was trained. Leaderboard ↗ActiveNot confirmed % resolved (pass@1)
SWE-Lancer OpenAI · 2025 CodingWhether frontier models can complete real paid freelance software jobs, both coding and technical-management tasks, well enough to earn the payouts. ActiveNot confirmed Dollars earned (and % of tasks resolved)
SWE-PolyBench Rashid, Bock, Zhuang et al. · 2025 CodingWhether a coding agent can resolve repository-level tasks outside Python, across Java, JavaScript, TypeScript and Python. ActiveNot confirmed % resolved, plus syntax-tree-based retrieval and file-localisation metrics
SWE-rebench Badertdinov, Golubev, Nekrashevich et al. · 2025 CodingIssue resolution on a continuously refreshed, decontaminated pool of Python software-engineering tasks mined automatically from open-source repositories. Leaderboard ↗ActiveNot confirmed % resolved (pass@1)
Terminal-Bench Stanford and the Laude Institute · 2026 Coding· Full guideWhether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal. Leaderboard ↗Active83.4% (v2.1) Codex (GPT-5.5) Pass/fail, graded by verification scripts in the agent's Docker environment (pass@1) · read 2026-06
AgentBench Tsinghua University · 2023 Agents & tool useHow well an LLM acts as an autonomous agent in multi-turn, open-ended decision-making across diverse interactive environments. Leaderboard ↗ActiveNot confirmed Per-environment success aggregated into an overall score
AndroidWorld Rawles, Clinckemaillie, Chang et al. · 2024 Agents & tool useWhether an agent can operate a real Android phone to finish tasks across everyday apps. Leaderboard ↗ActiveNot confirmed Programmatic reward from the device end state
AutomationBench Zapier · 2026 Agents & tool useWhether an agent can run a realistic business workflow end to end across several apps: discovering the right API endpoints itself, following a policy document, and writing correct data into every system it touches. Leaderboard ↗Active26.2% Claude Opus 5 (max effort) task_completed_correctly: strict pass/fail where every scored end-state assertion must pass, with partial credit reported only as a diagnostic · read 2026-07
BFCL UC Berkeley Gorilla team · 2024 Agents & tool use· Full guideWhether a model calls functions and APIs correctly: picking the right function, filling parameters with valid types, and refusing to invent functions that were not offered. Leaderboard ↗ActiveNot confirmed Abstract-syntax-tree match against a reference call, plus executable checks; overall score is the unweighted mean of subcategories
BrowseComp OpenAI · 2025 Agents & tool useWhether a browsing agent can persistently navigate the open web to locate a single hard-to-find, entangled fact. Active51.5% OpenAI Deep Research (launch paper) Accuracy via model-graded semantic equivalence to the reference answer · read 2025-04
DeepSearchQA Google DeepMind · 2026 Agents & tool useWhether a deep-research agent can plan and execute a long chain of web searches to return an exhaustive, de-duplicated answer list rather than a single fact. Leaderboard ↗ActiveNot confirmed Accuracy against each task's objectively verifiable exhaustive answer set
GAIA Meta AI and Hugging Face · 2023 Agents & tool use· Full guideWhether an AI assistant can answer real-world questions that require multi-step reasoning, multiple modalities, web browsing and general tool use. Leaderboard ↗Active~75% HAL agent (Claude Sonnet 4.5) Exact-match accuracy against an unambiguous answer · read 2026-06
LoCoMo Maharana, Lee, Tulyakov et al. · 2024 Agents & tool useWhether an agent remembers and reasons over a conversation that spans months, rather than a single session. ActiveNot confirmed Question-answering accuracy, event summarisation and multi-modal dialogue generation
Mind2Web 2 Gou, Huang, Ning et al. · 2025 Agents & tool useWhether an agentic search or deep-research system can browse the live web and return a correct, citation-backed answer to a long-horizon question. Leaderboard ↗ActiveNot confirmed Agent-as-a-Judge rubric scoring of answer correctness and citation support
MLE-bench OpenAI · 2024 Agents & tool useWhether an AI agent can do end-to-end machine-learning engineering (data prep, training, experimentation, submission) at the level of human Kaggle competitors. Leaderboard ↗Active16.9% (paper baseline) o1-preview with AIDE scaffolding Medal rate (fraction of competitions reaching bronze/silver/gold thresholds) · read 2024-10
OSWorld XLANG Lab, University of Hong Kong · 2024 Agents & tool useWhether a multimodal agent can operate a real computer (desktop apps, file I/O, multi-app workflows) to complete open-ended tasks in a live virtual machine. Leaderboard ↗SaturatedNot confirmed Execution-based success rate via per-task verification scripts that inspect machine state
OSWorld 2.0 XLANG Lab, University of Hong Kong, with collaborators at Columbia, UCSB, UCSD, Mila, Ohio State and others · 2026 Agents & tool use· Full guideWhether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state. Leaderboard ↗Active20.6% Claude Opus 4.8 (max thinking) Binary completion at a 500-step cap, reported with a weighted-checkpoint partial score · read 2026-06
tau-bench Sierra · 2024 Agents & tool use· Full guideWhether a tool-using agent can reliably complete customer-service tasks over multi-turn conversations with a simulated user while obeying domain policies. ActiveNot confirmed pass^k: the probability an agent succeeds across all k independent trials (reliability, not just average success)
TheAgentCompany Xu, Song, Li et al. · 2024 Agents & tool useWhether an agent can do real knowledge work inside a simulated software company: browsing, coding, using internal tools, and messaging simulated colleagues. Leaderboard ↗ActiveNot confirmed Full and partial task completion, scored by checkpoint
Vending-Bench Andon Labs · 2025 Agents & tool useWhether an agent stays coherent over a very long horizon, by running a simulated vending-machine business: stock, orders, pricing and daily fees. ActiveNot confirmed Net worth and units sold at the end of the run
VisualWebArena Carnegie Mellon University · 2024 Agents & tool useWhether a multimodal agent can complete visually grounded web tasks that require interpreting images and page layout, not just text. Leaderboard ↗ActiveNot confirmed Functional success rate via execution-based evaluation
WebArena Carnegie Mellon University · 2023 Agents & tool useWhether an autonomous agent can complete long-horizon, realistic web tasks (navigation, forms, multi-step workflows) in fully functional self-hosted websites. Leaderboard ↗ActiveNot confirmed Functional success rate via execution-based reward checking the end state
WebVoyager He, Yao, Ma et al. · 2024 Agents & tool useWhether a multimodal web agent can complete a user instruction end to end on real, live websites rather than a simulator or a static snapshot. SaturatedNot confirmed Task success rate, judged automatically from screenshots and responses
Windows Agent Arena Bonatti, Zhao, Bonacci et al. · 2024 Agents & tool useWhether a multimodal agent can operate a full Windows desktop across the applications people actually use at work. ActiveNot confirmed Task success rate from the OS end state
AGIEval Zhong, Cui, Guo et al. · 2023 ReasoningHuman-centric reasoning, using questions from real standardised exams taken by people rather than synthetic datasets. SaturatedNot confirmed Accuracy
ARC-AGI-1 Francois Chollet · 2019 ReasoningWhether a system can infer the abstract rule of a novel visual grid puzzle from a few examples and apply it to a new input. Leaderboard ↗Saturated97.5% (public eval) Claude Opus 5 and GPT-5.6 Sol pass@2 exact-grid-match accuracy · read 2026-07
ARC-AGI-2 ARC Prize Foundation · 2025 Reasoning· Full guideThe same fluid-intelligence test as ARC-AGI-1, but with harder, contamination-resistant tasks that stay easy for humans yet very hard for AI. Leaderboard ↗Saturated92.5% (semi-private) GPT-5.6 Sol (max effort) pass@2 exact-grid-match accuracy, reported with a cost-per-task efficiency metric · read 2026-07
ARC-AGI-3 ARC Prize Foundation · 2026 Reasoning· Full guideWhether an agent dropped into an unfamiliar interactive environment with no instructions, stated goal or rules can work out what to do by acting, build a usable world model, and keep learning across levels. Leaderboard ↗Active30.2% (public demo) Claude Opus 5 (high effort) Games beaten at or above human-level action efficiency, measuring skill-acquisition efficiency rather than one-shot accuracy · read 2026-07
BIG-Bench Hard Suzgun et al. · 2022 ReasoningA suite of multi-step reasoning tasks (logic, arithmetic, algorithmic, commonsense) on which pre-2022 models trailed average human raters. SaturatedNot confirmed Per-task accuracy averaged across the 23 tasks
DROP Dua, Wang, Dasigi et al. · 2019 ReasoningReading comprehension that requires discrete operations over a passage: resolving references then adding, counting or sorting. SaturatedNot confirmed F1 and exact match
EnigmaEval Scale AI · 2025 ReasoningLong multimodal puzzle solving: finding hidden connections between unrelated pieces of information and chaining many deductive steps. Leaderboard ↗ActiveNot confirmed Exact-match accuracy on the final puzzle answer
GPQA Diamond Rein et al. · 2023 Reasoning· Full guideGraduate and PhD-level multiple-choice scientific reasoning in biology, physics and chemistry, on questions designed to be unanswerable by quick web search. Leaderboard ↗Saturated~94% Gemini 3.1 Pro Preview Multiple-choice accuracy (random baseline 25%, PhD-expert baseline about 70%) · read 2026-02
Humanity's Last Exam Center for AI Safety (CAIS) and Scale AI · 2025 Reasoning· Full guideFrontier, closed-ended expert knowledge and reasoning across more than 100 academic disciplines at the limit of human expertise. Leaderboard ↗Active53.3% Claude Fable 5 (Max Effort) Accuracy (exact match / multiple-choice), often reported with a calibration metric · read 2026-06
MuSR Sprague, Ye, Durrett et al. · 2023 ReasoningMultistep commonsense reasoning embedded in long natural-language narratives such as murder mysteries, object placement and team allocation. Leaderboard ↗ActiveNot confirmed Multiple-choice accuracy
ZebraLogic Lin, Le Bras, Richardson et al. · 2025 ReasoningLogical deduction under hard constraints, using logic grid puzzles generated from constraint-satisfaction problems. ActiveNot confirmed Puzzle-level accuracy (all cells correct)
AIME 2025 Mathematical Association of America; adopted as an LLM eval by the community · 2025 MathematicsOlympiad-track competition mathematics at the level of the American Invitational Mathematics Examination, used as a high-difficulty LLM eval. Leaderboard ↗Saturated100% Multiple frontier reasoning models Exact-match accuracy, usually pass@1 averaged over samples · read 2026-06
FrontierMath Epoch AI · 2024 Mathematics· Full guideResearch-level original mathematics requiring hours to days of expert effort, across number theory, analysis, algebraic geometry and more. Leaderboard ↗Saturated87% (Tiers 1-3) Claude Fable 5 Accuracy (fraction with a correct, automatically verifiable final answer) · read 2026-06
GSM8K OpenAI · 2021 MathematicsMulti-step grade-school arithmetic word-problem reasoning. Leaderboard ↗Saturated~99.6% Frontier models broadly Exact-match accuracy on the final numeric answer · read 2026-05
MATH Hendrycks et al. · 2021 MathematicsStep-by-step solving of high-school competition mathematics across algebra, geometry, number theory, probability and precalculus. Leaderboard ↗Saturated~99% (MATH-500) GPT-5 Exact-match accuracy on the final boxed answer · read 2026-04
MathArena ETH Zurich · 2025 MathematicsMathematical reasoning and proof-writing on freshly released competition problems, evaluated before they can enter training data. Leaderboard ↗Active81.1% (aggregate) GPT-5.5 (xhigh) Per-competition accuracy and an aggregate expected-performance score · read 2026-04
miniF2F Zheng, Han and Polu · 2021 MathematicsFormal theorem proving on Olympiad-level mathematics, as a shared benchmark across proof assistants. SaturatedNot confirmed % of statements formally proved
Omni-MATH Gao, Song, Cai et al. · 2024 MathematicsOlympiad-level mathematical reasoning across a broad range of subdomains and difficulty levels. Leaderboard ↗ActiveNot confirmed Accuracy, scored with an LLM-based verifier (Omni-Judge)
PutnamBench Tsoukalas, Lee, Jennings et al. · 2024 MathematicsWhether a neural theorem prover can produce a formal, machine-checked proof of an undergraduate competition problem. Leaderboard ↗ActiveNot confirmed % of theorems formally proved and machine-verified
HellaSwag Zellers, Holtzman, Bisk et al. · 2019 KnowledgeCommonsense sentence completion: picking the plausible continuation of an everyday scenario. RetiredNot confirmed Accuracy
MMLU Hendrycks et al. · 2021 Knowledge· Full guideBroad academic and professional knowledge across 57 subjects via four-choice multiple-choice questions. Leaderboard ↗Saturated~93% Qwen3.7 Max Accuracy · read 2026-06
MMLU-Pro TIGER-Lab · 2024 Knowledge· Full guideHarder multi-task reasoning and knowledge designed to de-saturate MMLU and reward deliberate reasoning over recall. Leaderboard ↗Active~90% Gemini 3 Pro Preview Accuracy · read 2026-06
MMLU-Redux Gema et al. · 2024 KnowledgeA re-annotated, error-corrected subset of MMLU used to measure true knowledge accuracy without the original's label noise. Leaderboard ↗ActiveNot confirmed Accuracy on cleaned labels
SimpleQA OpenAI · 2024 KnowledgeShort-form parametric factuality: whether a model answers single-answer fact-seeking questions correctly and abstains when unsure. Leaderboard ↗ActiveNot confirmed Accuracy, plus correct-given-attempted and an F-score balancing attempts against accuracy
SuperGPQA M-A-P Team · 2025 KnowledgeGraduate-level knowledge and reasoning across 285 disciplines, including the applied and service fields that mainstream benchmarks ignore. ActiveNot confirmed Accuracy
TriviaQA Joshi, Choi, Weld et al. · 2017 KnowledgeFactual recall and reading comprehension over trivia questions with evidence documents. RetiredNot confirmed Exact match and F1
Global-MMLU Singh, Romanou, Fourrier et al. · 2024 Instructions & languagesMultilingual academic knowledge, separating questions that are culturally neutral from those requiring culture-specific knowledge. ActiveNot confirmed Accuracy
IFEval Zhou, Lu, Mishra et al. · 2023 Instructions & languagesWhether a model obeys instructions that can be checked by a program, such as a word count, a required keyword, or a forbidden format. SaturatedNot confirmed Strict and loose instruction-following accuracy
INCLUDE Romanou, Foroutan, Sotnikova et al. · 2024 Instructions & languagesMultilingual understanding built from local exam material, so the questions test regional knowledge rather than translated Western content. ActiveNot confirmed Accuracy
MGSM Shi, Suzgun, Freitag et al. · 2022 Instructions & languagesGrade-school math word problems solved via chain-of-thought reasoning in ten languages. SaturatedNot confirmed Accuracy
Multi-IF He, Jin, Wang et al. · 2024 Instructions & languagesWhether a model keeps following instructions across multiple turns and in languages other than English. ActiveNot confirmed Instruction-following accuracy per turn
BABILong Kuratov, Bulatov, Anokhin et al. · 2024 Long contextReasoning over facts deliberately scattered through an extremely long document, not just retrieving one of them. ActiveNot confirmed Accuracy by context length
HELMET Yen, Gao, Hou et al. · 2024 Long contextLong-context ability across a wide spread of realistic downstream applications rather than one synthetic retrieval task. Leaderboard ↗ActiveNot confirmed Per-category task metrics, reported across context lengths
LOFT Lee, Chen, Dai et al. · 2024 Long contextWhether a long-context model can replace a retrieval pipeline outright: doing retrieval, RAG and SQL-style tasks natively from context. ActiveNot confirmed Task-specific accuracy compared against specialised retrieval pipelines
LongBench Tsinghua University · 2023 Long contextComprehensive long-context understanding across realistic tasks (QA, summarization, few-shot, code, synthetic) in English and Chinese. Leaderboard ↗Active57.7% (v2, with reasoning) o1-preview v1: per-task automatic metrics. v2: multiple-choice accuracy · read 2024-12
MRCR Google DeepMind (Michelangelo); open-source variant by OpenAI · 2024 Long contextWhether a model can distinguish and retrieve the correct one among multiple near-identical requests buried in a long multi-turn conversation. ActiveNot confirmed Similarity of the model’s output to the target instance, gated by a required answer-prefix
Needle-in-a-Haystack Greg Kamradt · 2023 Long contextWhether a model can recall a single planted fact (the needle) inserted at varying depths within a long context (the haystack). SaturatedNot confirmed Retrieval accuracy at each depth and length cell
NoLiMa Adobe Research and LMU Munich · 2025 Long contextLong-context retrieval and reasoning when the question and the target fact share minimal literal word overlap, forcing latent association rather than keyword matching. Leaderboard ↗ActiveNot confirmed Accuracy at each length, relative to the model's short-context baseline
RULER NVIDIA · 2024 Long context· Full guideThe real effective context length of a model by testing retrieval, multi-hop tracing, aggregation and QA at increasing sequence lengths. Leaderboard ↗ActiveNot confirmed Weighted-average accuracy across tasks and lengths; effective length is the longest length still above threshold
ChartQA Masry, Long, Tan et al. · 2022 MultimodalQuestion answering over charts that requires both reading visual features and doing arithmetic on them. SaturatedNot confirmed Relaxed accuracy (numeric answers within a tolerance)
CharXiv Wang, Xia, He et al. · 2024 MultimodalChart understanding on real, messy scientific figures rather than clean template-generated charts. Leaderboard ↗ActiveNot confirmed Accuracy, split into descriptive and reasoning questions
DocVQA Mathew, Karatzas and Jawahar · 2020 MultimodalQuestion answering over scanned document images, where layout and structure carry the meaning. SaturatedNot confirmed ANLS (average normalised Levenshtein similarity)
MathVista Lu et al. · 2023 MultimodalMathematical and quantitative reasoning grounded in visual contexts such as figures, charts, geometry and scientific diagrams. Leaderboard ↗Active~91% (testmini) Seed 2.1 Pro Accuracy · read 2026-06
MMBench Liu, Duan, Zhang et al. · 2023 MultimodalFine-grained vision-language ability across a structured taxonomy of perception and reasoning skills. Leaderboard ↗SaturatedNot confirmed Accuracy under circular evaluation
MMMU MMMU team · 2023 Multimodal· Full guideCollege-level multimodal understanding and reasoning over images, diagrams, charts and text across many disciplines. Leaderboard ↗Active~86% Qwen3.6 Plus Accuracy · read 2026-06
MMMU-Pro MMMU team · 2024 MultimodalA harder, contamination-resistant version of MMMU that forces genuine visual reasoning rather than text-only shortcuts. Leaderboard ↗Active~84% Gemini 3.5 Flash Accuracy · read 2026-06
MMStar Chen, Li, Dong et al. · 2024 MultimodalGenuinely vision-dependent multimodal ability, on samples selected so the answer cannot be inferred from the text alone. ActiveNot confirmed Accuracy, reported alongside a multimodal-gain and multimodal-leakage measure
Video-MME MME-Benchmarks team · 2024 MultimodalComprehensive video understanding by multimodal LLMs across short, medium and long clips. Leaderboard ↗Active~89% Seed 2.1 Pro Accuracy (tested with and without subtitles) · read 2026-06
FinanceBench Patronus AI · 2023 Professional domainsOpen-book financial question answering over real public-company filings, with the supporting evidence required. ActiveNot confirmed Answer correctness against the evidence, human reviewed
HealthBench OpenAI · 2025 Professional domains· Full guideOpen-ended clinical conversation quality and safety, graded against rubrics written by practising physicians. ActiveNot confirmed Rubric score, graded by a model grader against physician-written criteria
LegalBench Guha, Nyarko, Ho et al. · 2023 Professional domains· Full guideLegal reasoning across the specific skills lawyers actually use, as defined by legal professionals. Leaderboard ↗ActiveNot confirmed Per-task accuracy, aggregated by reasoning type
MedHELM Stanford CRFM · 2025 Professional domainsClinical ability across the breadth of real medical work, on a clinician-validated taxonomy rather than exam questions. Leaderboard ↗ActiveNot confirmed Per-task clinical metrics plus head-to-head win rates
MedQA Jin, Pan, Oufattole et al. · 2020 Professional domainsMedical knowledge, using real questions from professional medical board examinations. SaturatedNot confirmed Accuracy
AlpacaEval 2 (Length-Controlled) Dubois, Galambosi, Liang et al. · 2024 Human preferenceInstruction-following quality judged by an LLM, with a regression correction for the judge’s bias toward longer answers. Leaderboard ↗SaturatedNot confirmed Length-controlled win rate
Arena-Hard-Auto Li, Chiang, Frick et al. · 2024 Human preferenceHuman-preference-aligned quality on hard open-ended prompts, scored automatically instead of by live human voting. ActiveNot confirmed Win rate against a baseline model, judged by an LLM
Artificial Analysis Coding Agent Index Artificial Analysis · 2026 Human preferenceOverall coding-agent capability as a single number, scoring the full stack (a specific model plus its harness and settings) rather than a model in isolation. Leaderboard ↗ActiveNot confirmed Simple average of the component benchmark scores, with every task equally weighted
Artificial Analysis Intelligence Index Artificial Analysis · 2024 Human preferenceA composite index of overall model intelligence aggregating performance across reasoning, coding, knowledge, science and agentic tasks. Leaderboard ↗Active~60 (index) Claude Fable 5 Composite index score (0 to 100 aggregate) · read 2026-06
Copilot Arena Chi, Chen, Angelopoulos et al. · 2025 Human preferenceWhich coding model developers actually prefer, collected from paired completions inside a real editor rather than a chat window. ActiveNot confirmed Elo-style ranking from in-editor pairwise preferences
Epoch Capabilities Index Epoch AI · 2025 Human preference· Full guideOverall model capability on one continuous scale, stitched together from many benchmarks of differing difficulty. Leaderboard ↗ActiveNot confirmed ECI score on the anchored scale
GDPval OpenAI, with an agentic re-run by Artificial Analysis as GDPval-AA · 2025 Human preferenceWhether a model can produce the actual deliverables of skilled professional work (documents, slides, spreadsheets, diagrams) well enough to stand against an industry expert's version. Leaderboard ↗Active1861 Elo (GDPval-AA v2) Claude Opus 5 (adaptive reasoning, max effort) Blind pairwise comparison of two anonymised outputs on the same task, aggregated into an Elo rating; the v2 scale anchors human expert deliverables at 1000 · read 2026-07
HELM Stanford CRFM · 2022 Human preferenceMulti-metric holistic evaluation across many scenarios, reporting accuracy alongside calibration, robustness, fairness, bias, toxicity and efficiency. Leaderboard ↗ActiveNot confirmed Multi-metric (per-metric scores across scenarios; no single headline number)
LiveBench White, Dooley, Roberts et al. · 2024 Human preference· Full guideBroad capability across six categories at once (math, coding, reasoning, data analysis, instruction following, language), on questions refreshed monthly. Leaderboard ↗ActiveNot confirmed Objective automatic scoring against ground truth, averaged across categories
LMArena Arena · 2023 Human preference· Full guideCrowdsourced human preference between two anonymized model responses, aggregated into a relative ranking, not an objective capability. Leaderboard ↗Active~1510 Elo Claude Opus 4.8 Elo / Bradley-Terry pairwise rating (an Arena Score) · read 2026-06
METR Time Horizon METR · 2025 Human preference· Full guideModel capability expressed in human time: the length of task, measured by how long humans take, that a model completes with 50% success. ActiveNot confirmed 50%-task-completion time horizon, in minutes or hours
MT-Bench LMSYS · 2023 Human preferenceInstruction-following and conversational quality on multi-turn prompts, scored automatically by a strong LLM judge. Leaderboard ↗SaturatedNot confirmed LLM-as-judge score (1 to 10 scale, averaged)
HaluEval Li et al. · 2023 TruthfulnessA model's ability to recognize hallucinated content across question answering, knowledge-grounded dialogue and summarization. ActiveNot confirmed Hallucination-recognition accuracy (faithful vs hallucinated)
TruthfulQA Lin, Hilton, Evans · 2021 Truthfulness· Full guideWhether a model avoids repeating common human misconceptions when answering questions, rather than imitating popular falsehoods. ActiveNot confirmed % truthful (and % truthful-and-informative)
Vectara Hallucination Leaderboard Vectara · 2023 TruthfulnessHow often a model introduces unsupported content when summarizing a provided source document, i.e. faithfulness in closed-book summarization. Leaderboard ↗Active1.8% (lower is better) antgroup/finix-s1-32b Hallucination rate (% of summaries judged unfaithful; lower is better) · read 2026-05
AgentHarm Andriushchenko, Souly, Dziemian et al. · 2024 SecurityWhether a tool-using agent refuses explicitly malicious multi-step tasks, and whether it stays capable enough to complete them once jailbroken. ActiveNot confirmed Refusal rate and post-jailbreak task-completion rate
Cybench Zhang, Perry, Dulepet et al. · 2024 SecurityWhether an agent can autonomously solve professional capture-the-flag security tasks: finding a vulnerability and executing an exploit. Leaderboard ↗ActiveNot confirmed % of tasks and subtasks solved unassisted
CyberSecEval 3 Meta · 2024 SecurityCybersecurity risk across eight areas, split between risk to third parties and risk to the developers and users of an application. ActiveNot confirmed Per-risk pass and failure rates, measured with and without guardrails
WMDP Li, Pan, Gopal et al. · 2024 SecurityProxy knowledge of hazardous biosecurity, cybersecurity and chemical-security material. Leaderboard ↗ActiveNot confirmed Accuracy, where lower is the desired direction after unlearning

Top scores are representative snapshots as of July 27, 2026, not live readings: leaderboards move constantly, and many figures come from each benchmark's own board (the Leaderboard link in each row). Where no clean current top score could be confirmed from a primary source, the cell reads "Not confirmed." Confirm at the source before quoting a number.

Read the full guide to any benchmark

A row in the tables above tells you what a benchmark is. These 27 guides tell you whether to believe it: how the scoring actually works, the version history and saturation trajectory, who runs it and what stake they have, and the specific thing the headline number hides. Start with whichever benchmark a vendor just quoted at you.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

How to read a benchmark score

A headline score is only as good as the benchmark behind it. Four failure modes decide whether a number means anything, and the strongest benchmarks are the ones that resist all four. The scorecard below rates the most-quoted benchmarks on each concern, where higher always means a less trustworthy score. For the full argument and the receipts, read are AI benchmarks reliable and our breakdown of benchmark contamination.

Benchmark Trust ScorecardEach major AI benchmark rated Low, Medium, or High on four concern axes (contamination, saturation, gameability, and real-world gap), where higher always means a less trustworthy score. Full ratings are in the data table below.Low concernMediumHigh concernBenchmarkContaminationSaturationGameabilityReal-world gapMMLUknowledge MCQHighHighHighHighSWE-bench Verifiedreal GitHub fixesHighHighHighMediumLMArenahuman preferenceMediumLowHighHighGPQA DiamondPhD-level scienceMediumHighMediumMediumARC-AGI v2abstract reasoningLowHighMediumHighARC-AGI-3interactive reasoningLowLowMediumHighFrontierMathresearch mathMediumHighLowMediumTerminal-Bench 2terminal tasksLowMediumLowLowFrontier-Benchsenior engineeringLowLowMediumLow
Benchmark Trust Scorecard. Higher concern means a less trustworthy score.
BenchmarkContaminationSaturationGameabilityReal-world gap
MMLU (knowledge MCQ)HighHighHighHigh
SWE-bench Verified (real GitHub fixes)HighHighHighMedium
LMArena (human preference)MediumLowHighHigh
GPQA Diamond (PhD-level science)MediumHighMediumMedium
ARC-AGI v2 (abstract reasoning)LowHighMediumHigh
ARC-AGI-3 (interactive reasoning)LowLowMediumHigh
FrontierMath (research math)MediumHighLowMedium
Terminal-Bench 2 (terminal tasks)LowMediumLowLow
Frontier-Bench (senior engineering)LowLowMediumLow
The most-quoted benchmarks rated on four concerns, where higher means a less trustworthy score. The newest execution-graded designs (Frontier-Bench, Terminal-Bench 2) and the interactive ARC-AGI-3 sit greenest; the oldest public sets (MMLU) sit reddest. ARC-AGI v2 shows how fast a row turns: contamination-resistant by design, but saturated inside a year.Source: Capital & Compute benchmark trust scorecard, July 27, 2026

Which benchmark should you actually watch?

Pick by the decision you are making, not by whichever number a vendor leads with.

  • Choosing a coding model: read a completion benchmark and a quality benchmark together. DeepSWE measures whether an agent finishes the task; FrontierCode measures whether you would merge its code, and the same model can ace one while failing the other. Add Terminal-Bench for agentic, run-it-in-a-real-terminal work, and Frontier-Bench for senior-level feature and debugging tasks, where the best models still sit near 34%. Treat CursorBench with more caution: it is Anysphere's private suite, so nobody outside the company can reproduce a number from it.
  • Judging raw reasoning: the ARC ladder moved fast in 2026. ARC-AGI-1 is finished at 97.5% and ARC-AGI-2 went from 54% in December 2025 to an ARC-Prize-verified 92.5% by July, so neither separates the frontier any more. ARC-AGI-3 is now the widest gap on this page: humans scored 100% at its March 2026 launch against 0.51% for frontier models, and the verified top is still only 30.2%. Humanity’s Last Exam remains the unsaturated expert-breadth test at about 53%. Movement on those two is real progress, not noise.
  • Comparing general capability: a human-preference ranking like LMArena is the closest to "which feels better to use," but it rewards style as much as substance, so pair it with a composite index and a hard reasoning score.
  • Ranking agents, not chatbots: the same team now runs Agent Arena, which scores agents from real usage using causal tracing instead of votes, a method built to resist the style-gaming that dogs preference boards. The Artificial Analysis Coding Agent Index attacks the same problem from the other side, scoring a model and its harness together as one stack, which is what you actually deploy. See cost per task for what those scores cost to achieve.
  • Ignore the saturated ones: MMLU, GSM8K, MATH and HumanEval are quoted out of habit. When every frontier model scores above 95%, the benchmark is measuring the ceiling, not the model.

For how these benchmarks rank the current crop of agents, see how the 2026 coding-agent benchmarks actually rank, the coding agents that wrap these models, and the value leaderboard for points-per-dollar.

Why this is a snapshot, not a live feed

Benchmark scores are the most volatile, most gamed layer of model marketing. Leaderboards re-rank weekly, labs report only the tests they win, and the same benchmark can be run under different scaffolds that move the number by ten points or more. This directory is therefore a representative map as of July 27, 2026, built to show what each benchmark means and ground it to its own source, not to quote a score you can hold anyone to. The durable value is the left of the table (what it measures, who built it, whether it is still meaningful); the top-score column is a pointer to the live leaderboard, where you should always confirm before quoting a number.

Frequently asked questions

What is an AI benchmark?
An AI benchmark is a standardized test made of a fixed dataset, a task specification, and a scoring metric, used to measure and compare how well AI models perform a specific skill such as reasoning, coding, math, or knowledge. Running many models through the same test produces a single comparable score, which is what model leaderboards rank.
What are the main AI benchmarks in 2026?
For coding, SWE-bench Verified, DeepSWE, and Terminal-Bench. For reasoning, GPQA Diamond, ARC-AGI-2, and Humanity’s Last Exam. For math, FrontierMath and AIME. For broad knowledge, MMLU-Pro. For multimodal, MMMU. And for overall human preference, the LMArena Elo ranking. Many older benchmarks like MMLU, GSM8K, and HumanEval are now saturated and quoted mainly out of habit.
Are AI benchmarks reliable?
Partly. A benchmark is reliable only as far as its score reflects real capability, and four things erode that: contamination (test data leaking into training), saturation (top models bunched near the ceiling), gameability (a score inflated without real skill), and vendor cherry-picking (a lab reporting only the benchmarks it wins). The most trustworthy benchmarks are contamination-resistant, unsaturated, and run by an independent party. For the full breakdown, see our guide on whether AI benchmarks are reliable.
What does it mean when a benchmark is saturated?
A benchmark is saturated when the strongest models all score near its ceiling, so the differences between them are within noise and the benchmark no longer separates a better model from a worse one. MMLU, GSM8K, MATH, and HumanEval are all saturated in 2026, with top models above 95%, which is why the field keeps building harder replacements like MMLU-Pro and ARC-AGI-2.
What is the difference between SWE-bench and SWE-bench Verified?
SWE-bench is the original 2,294-task set of real GitHub issues. SWE-bench Verified is a 500-task subset that OpenAI and the SWE-bench authors hand-checked in 2024 to remove broken tests and unsolvable issues, so it is the cleaner, more-quoted version. By 2026 even Verified is treated as saturated and contamination-prone, and OpenAI now recommends the harder SWE-bench Pro instead.
Which benchmarks evaluate AI agent reliability in 2026?
The main agent-reliability benchmarks in 2026 are tau-bench, which measures whether a tool-using agent completes multi-turn tasks consistently across repeated runs; Terminal-Bench for hard end-to-end command-line tasks; OSWorld 2.0 for computer-use agents on long-horizon desktop work; AutomationBench for cross-application business workflows graded on end state; GAIA and AgentBench for multi-step assistant tasks; and WebArena for long-horizon web tasks. Agent Arena complements these by ranking agents from real usage instead of a fixed test set.
What are the new AI benchmarks in 2026?
The benchmarks that appeared or replaced predecessors in 2026 are Frontier-Bench, 74 senior-level engineering tasks from the Terminal-Bench team; CursorBench, Anysphere private suite mined from real Cursor sessions; OSWorld 2.0, 108 long-horizon computer-use tasks that replace the now-solved OSWorld; AutomationBench, 600 or more cross-app business workflows from Zapier; DeepSearchQA, 900 exhaustive-answer research prompts from Google DeepMind; ARC-AGI-3, the first fully interactive ARC where frontier models scored 0.51% at launch; and the Artificial Analysis Coding Agent Index, which scores whole agent stacks rather than bare models. They exist because the 2024 generation of benchmarks saturated.
Which AI benchmark matters most for coding?
There is no single one, because they measure different things. SWE-bench Verified and DeepSWE measure whether an agent can resolve a real issue; Terminal-Bench measures whether it can operate a real terminal end to end; FrontierCode measures whether the code is clean enough to merge; LiveCodeBench measures contamination-free competitive programming. For shipping production code, read a completion benchmark and a quality benchmark together rather than trusting one number.
What is benchmark contamination?
Contamination is when a benchmark’s questions or answers leak into a model’s training data, so the model can recall the answer instead of reasoning it out. It inflates scores without reflecting real capability and is the main reason public, static benchmarks decay over time. The defenses are private or held-out test sets, time-stamped problems released after a model’s cutoff, and freshly generated tasks.
How many AI benchmarks are there?
There is no fixed number, because new benchmarks are created continuously as older ones saturate. This directory tracks 116 of the benchmarks frontier labs actually report, across 12 categories, 28 of which are already saturated. In practice roughly a dozen benchmarks carry most of the signal in any given model launch.
What is the hardest AI benchmark in 2026?
By the size of the remaining human-model gap it is ARC-AGI-3, where humans scored 100% at the March 2026 launch against 0.51% for frontier models, and the ARC-Prize-verified top is still around 30%. Humanity’s Last Exam is the hardest knowledge benchmark at roughly 53%, and OSWorld 2.0 is the hardest agent benchmark at about 20% full task completion.
Which AI benchmarks are contamination resistant?
The ones that do not rely on a fixed public test set. Four designs work: private or held-out splits, as in ARC-AGI-2 and GAIA; time-stamped problems released after a model training cutoff, as in LiveCodeBench and SWE-bench-Live; continuously rotated questions, as in LiveBench; and programmatically generated tasks, as in ZebraLogic and RULER. Formal proof benchmarks such as PutnamBench are also hard to fake, because a proof checker either accepts the proof or it does not.
What is the difference between AI benchmarks and MLPerf?
They measure different layers. Model benchmarks such as MMLU, SWE-bench and Terminal-Bench measure what a model can do, scored as accuracy or task completion. Systems benchmarks such as MLPerf measure how fast and how efficiently hardware runs a fixed workload, scored as throughput, latency and energy. A model benchmark helps you pick a model; a systems benchmark tells you what serving it costs.
How long does an AI benchmark stay useful?
Between one and three years, and the window is shrinking. The 2021 cohort, including HumanEval, GSM8K and MMLU, stayed useful for roughly three years. Benchmarks released in 2024 have lasted about two. ARC-AGI-2, released in 2025, went from 54% to an ARC-Prize-verified 92.5% within a year. A benchmark score is therefore only meaningful with a date attached.
What is GPQA Diamond, and why is it hard?
GPQA Diamond is a 198-question set of graduate and PhD-level biology, physics, and chemistry questions written by domain experts to be Google-proof, meaning a non-expert with web access still cannot answer them quickly. It tests reasoning over retrieval. By 2026 top models exceed the roughly 70% human-expert baseline and sit in the low-to-mid 90s, so it is now largely saturated.

Sources

Each benchmark is grounded to its primary source: the original paper, the project repository, or the official leaderboard. Rows are re-verified in batches rather than all on one day, so each entry below carries the date it was last checked against its source (June to July 2026). Top scores are representative snapshots from those sources:

Machine-readable data: /ai-benchmarks.json. Benchmark reliability ratings are from our benchmark trust scorecard.

← All tools & trackers