TestsBenchmarks
What the testsWhat the benchmarksactually measure.
AI models are compared by giving them standard tests. We follow 207 of them. 63 feed our scores: 9 for writing, 27 for research and analysis, 29 for coding.207 benchmarks in plain English. 63 feed a use-case composite (writing 9, analysis 27, coding 29), weighted by how well they predict real work.
Used in our scoresIn a use-case composite
These tests feed the writing, research-and-analysis and coding scores you see on the front page. Some count for more than one area.
Benchmarks with use-case weight ≥ 0.5 that aren't saturated. Weights per use case shown on each row.
GDPval-AA v2.1
Counts toward:Weights: Writing ×0.5 · Research ×1GDPval-AA version 2.1 (Elo).
Higher is better · 27 models tested · Elo rating · higher is better · 27 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
1846 · Max
SWE-Bench Pro (public, v1)
Counts toward:Weights: Coding ×1Share of 731 long-horizon, multi-file tasks from copyleft open-source repos an agent resolves without breaking existing tests.
Higher is better · 45 models tested · % solved · higher is better · 45 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
89.9% · Max
AA-LCR
Counts toward:Weights: Research ×1Artificial Analysis long-context reasoning.
Higher is better · 43 models tested · % solved · higher is better · 43 models scored · the test's websiteofficial page ↗
Kimi K3
88.7% · Max
Terminal-Bench 4.0
Counts toward:Weights: Coding ×1Share of 66 hard, real terminal tasks (software, science, ML, ops, security, hardware) an agent completes end-to-end in a container.
Higher is better · 41 models tested · % solved · higher is better · 41 models scored · the test's websiteofficial page ↗
Claude Sonnet 5.5
66.2% · Max
LMArena Text - Creative Writing
Counts toward:Weights: Writing ×1Crowd-voted, style-controlled Elo on prompts classified as creative writing (stories, poems, scripts, original prose).
Higher is better · 28 models tested · Elo rating · higher is better · 28 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1522 · High
LMArena Text - Occupational: Writing, Literature & Language
Counts toward:Weights: Writing ×1Crowd-voted, style-controlled Elo on prompts from writing/editing, literature and language work (drafting, rewriting, editing, translation-style tasks).
Higher is better · 27 models tested · Elo rating · higher is better · 27 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1523 · High
LMArena Text - Longer Query
Counts toward:Weights: Writing ×0.5 · Research ×0.5Crowd-voted, style-controlled Elo on long user prompts (pasted documents, long briefs).
Higher is better · 25 models tested · Elo rating · higher is better · 25 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1544 · High
Artificial Analysis Coding Agent Index
Counts toward:Weights: Coding ×1Equal-weight composite of DeepSWE, Terminal-Bench 4.0 and SWE-Atlas-QnA, measured for model+agent pairs (Claude Code, Codex, etc.).
Higher is better · 21 models tested · Score · higher is better · 21 models scored · the test's websiteofficial page ↗
Claude Sonnet 5.5
68.4 · Max
DeepSWE
Counts toward:Weights: Coding ×0.9Pass@1 on 113 from-scratch, long-horizon engineering tasks across 91 repos and 5 languages, with hand-written behavioral verifiers.
Higher is better · 45 models tested · % solved · higher is better · 45 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
78.8% · Default
AA-Briefcase v1.1
Counts toward:Weights: Research ×0.9AA-Briefcase version 1.1 (Elo).
Higher is better · 27 models tested · Elo rating · higher is better · 27 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
1822 · Max
CursorBench
Counts toward:Weights: Coding ×0.9Share of ambiguous, multi-file tasks drawn from real Cursor sessions that an agent completes inside Cursor, per model and effort level.
Higher is better · 15 models tested · % solved · higher is better · 15 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
57.8% · Max
SWE-Bench Pro (private/commercial set)
Counts toward:Weights: Coding ×0.9Share of 276 tasks from private startup codebases an agent resolves; tests generalization to code models have never seen.
Higher is better · 8 models tested · % solved · higher is better · 8 models scored · the test's websiteofficial page ↗
Muse Spark 1.1
51.5% · Default
AA-Omniscience Index
Counts toward:Weights: Research ×0.8Knowledge-reliability index (-100..100): rewards correct answers, penalises wrong answers, abstaining is neutral; questions across business, health, law, humanities, science and software domains.
Higher is better · 43 models tested · Score · higher is better · 43 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
46.4 · Max
AA-Omniscience Hallucination Rate
Counts toward:Weights: Research ×0.8Of the questions a model did not get right, the share it answered wrongly instead of abstaining (lower = admits uncertainty instead of making things up).
Lower is better · 32 models tested · % solved · lower is better · 32 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
15.1% · High
Vals Legal Research Bench
Counts toward:Weights: Research ×0.8Agentic legal research tasks across US law areas, graded on correctness of the research answer.
Higher is better · 30 models tested · % solved · higher is better · 30 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
55.3% · Max
FrontierSWE
Counts toward:Weights: Coding ×0.8Mean@5 score on 34 very long (up to 20h) expert-level engineering and research-engineering tasks.
Higher is better · 13 models tested · % solved · higher is better · 13 models scored · the test's websiteofficial page ↗
Kimi K3
81.2% · Max
SWE-Bench Pro V2 (hard)
Counts toward:Weights: Coding ×0.8Hard subset of SWE-Bench Pro V2; still separates frontier agents better than the full set.
Higher is better · 11 models tested · % solved · higher is better · 11 models scored · the test's websiteofficial page ↗
Claude Opus 5
98.0% · Extra high
LLM Creative Story-Writing Benchmark (Lech Mazur)
Counts toward:Weights: Writing ×0.7Short stories to constrained briefs, compared pairwise by a panel of LLM evaluators; Thurstone comparison score centered at 0.
Higher is better · 44 models tested · Score · higher is better · 44 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
3.8 · High
EQ-Bench Creative Writing v3 (Elo)
Counts toward:Weights: Writing ×0.7LLM-judged (Claude Sonnet 4.6) pairwise Elo on 32 creative-writing prompts x3 with anti-slop/length controls.
Higher is better · 31 models tested · Elo rating · higher is better · 31 models scored · the test's websiteofficial page ↗
GPT-6 Astra
2173 · Default
FrontierCode
Counts toward:Weights: Coding ×0.7Measures whether a maintainer would actually merge the agent PR: correctness, tests, scope and code-quality, graded by an ensemble.
Higher is better · 25 models tested · % solved · higher is better · 25 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
54.6% · Medium
PRBench Legal (Scale)
Counts toward:Weights: Research ×0.7Expert-rubric graded professional legal reasoning questions (Scale SEAL).
Higher is better · 19 models tested · % solved · higher is better · 19 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
61.6% · Default
FrontierCode v1.1 (Extended)
Counts toward:Weights: Coding ×0.7Extended task set of Cognition's FrontierCode with the same correctness + mergeability grading.
Higher is better · 8 models tested · % solved · higher is better · 8 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
65.3% · Medium
Artificial Analysis Intelligence Index
Counts toward:Weights: Research ×0.6Composite of ~10 evals (HLE, GPQA, Terminal-Bench, tau2, SciCode, IFBench, etc.) run by Artificial Analysis per model and effort setting.
Higher is better · 61 models tested · Score · higher is better · 61 models scored · the test's websiteofficial page ↗
GPT-6 Astra
61.2 · Default
SWE Atlas - Codebase QnA
Counts toward:Weights: Coding ×0.6Rubric-graded accuracy answering deep technical questions about large real codebases using an agent harness.
Higher is better · 40 models tested · % solved · higher is better · 40 models scored · the test's websiteofficial page ↗
Claude Sonnet 5.5
66.9% · Max
SimpleQA Verified
Counts toward:Weights: Research ×0.6Short factual parametric-knowledge questions.
Higher is better · 38 models tested · % solved · higher is better · 38 models scored · the test's websiteofficial page ↗
GPT-6 Astra
75.6% · Max
LMArena Text - Occupational: Legal & Government
Counts toward:Weights: Research ×0.6Crowd-voted, style-controlled Elo on prompts from legal and government/public-sector work.
Higher is better · 29 models tested · Elo rating · higher is better · 29 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1537 · High
LMArena Code Arena (WebDev)
Counts toward:Weights: Coding ×0.6Crowd-voted Elo from head-to-head comparisons of models building web apps agentically in Code Arena.
Higher is better · 27 models tested · Elo rating · higher is better · 27 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
1818 · Max
LMArena Text - Instruction Following
Counts toward:Weights: Writing ×0.6Crowd-voted, style-controlled Elo on prompts with explicit constraints/instructions.
Higher is better · 25 models tested · Elo rating · higher is better · 25 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1529 · High
Vals Code Migration
Counts toward:Weights: Coding ×0.6How well an agent re-implements real programs in another language (CLI tools, COBOL to Java) with tests and code-quality checks.
Higher is better · 25 models tested · % solved · higher is better · 25 models scored · the test's websiteofficial page ↗
Claude Sonnet 5.5
69.8% · Max
LMArena Text - Expert
Counts toward:Weights: Research ×0.6Crowd-voted, style-controlled Elo on expert-level prompts (deep domain knowledge questions).
Higher is better · 24 models tested · Elo rating · higher is better · 24 models scored · the test's websiteofficial page ↗
Claude Fable 5
1549 · High
Vals Public Benefits Bench v1.1
Counts toward:Weights: Research ×0.6Can a model correctly help people navigate public-benefit (US SNAP) rules and eligibility - policy-text application.
Higher is better · 24 models tested · % solved · higher is better · 24 models scored · the test's websiteofficial page ↗
Claude Opus 5
76.9% · Max
SWE Atlas - Test Writing
Counts toward:Weights: Coding ×0.6Quality of unit/integration tests an agent writes for real codebases, graded against rubrics and mutation checks.
Higher is better · 23 models tested · % solved · higher is better · 23 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
67.0% · Extra high
Vals Vibe Code Bench
Counts toward:Weights: Coding ×0.6Share of from-scratch web applications an agent builds that pass end-to-end functional checks.
Higher is better · 23 models tested · % solved · higher is better · 23 models scored · the test's websiteofficial page ↗
Claude Sonnet 5.5
92.4% · Max
BrowseComp
Counts toward:Weights: Research ×0.6Hard-to-find web information retrieval by a browsing agent.
Higher is better · 22 models tested · % solved · higher is better · 22 models scored · the test's websiteofficial page ↗
GPT-6 Astra
91.5% · Default
Vals CorpFin v2
Counts toward:Weights: Research ×0.6Private QA over long credit agreements (long-document financial analysis).
Higher is better · 22 models tested · % solved · higher is better · 22 models scored · the test's websiteofficial page ↗
Claude Opus 5
73.2% · Default
ProgramBench (avg test pass rate)
Counts toward:Weights: Coding ×0.6Average share of hidden behavioral tests passed when re-implementing 200 programs from scratch; partial-credit view of ProgramBench.
Higher is better · 20 models tested · % solved · higher is better · 20 models scored · the test's websiteofficial page ↗
Claude Opus 5
74.7% · Extra high
SWE Atlas - Refactoring
Counts toward:Weights: Coding ×0.6Share of realistic large refactoring tasks an agent completes correctly in real repos.
Higher is better · 17 models tested · % solved · higher is better · 17 models scored · the test's websiteofficial page ↗
GPT-6 Astra
59.0% · Extra high
PRBench Finance (Scale)
Counts toward:Weights: Research ×0.6Expert-rubric graded professional finance reasoning questions (Scale SEAL).
Higher is better · 15 models tested · % solved · higher is better · 15 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
59.5% · Default
Harvey LAB-AA
Counts toward:Weights: Research ×0.6Harvey legal agent benchmark as run by Artificial Analysis (score).
Higher is better · 14 models tested · % solved · higher is better · 14 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
95.5% · Extra high
SWE-rebench
Counts toward:Weights: Coding ×0.6Resolve rate on fresh, decontaminated GitHub issues collected in a rolling time window, run with one standard scaffold.
Higher is better · 13 models tested · % solved · higher is better · 13 models scored · the test's websiteofficial page ↗
Claude Fable 5
64.5% · High
Terminal-Bench 2.0
Counts toward:Weights: Coding ×0.6Agentic tasks in a command-line environment (planning, iteration, tool use).
Higher is better · 12 models tested · % solved · higher is better · 12 models scored · the test's websiteofficial page ↗
GPT-5.5
82.7% · Extra high
GSO
Counts toward:Weights: Coding ×0.6Share of 102 software-optimization tasks where one attempt reaches at least 95% of the expert speedup while passing correctness tests.
Higher is better · 11 models tested · % solved · higher is better · 11 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
88.2% · Extra high
FrontierSWE v2
Counts toward:Weights: Coding ×0.6Proximal's 34 ultra-long-horizon engineering/research tasks (~20h per task), run in Proximal's harness at max effort.
Higher is better · 8 models tested · % solved · higher is better · 8 models scored · the test's websiteofficial page ↗
GPT-6 Astra
65.5% · Max
CursorBench 3.2.0
Counts toward:Weights: Coding ×0.6Earlier CursorBench release; not comparable to 4.0.
Higher is better · 6 models tested · % solved · higher is better · 6 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
73.4% · Max
Frontier-Bench v0.1
Counts toward:Weights: Coding ×0.6Agentic terminal-coding benchmark (74 tasks) run via Harbor or mini-SWE-agent; reported at Claude Opus 5 launch.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Claude Opus 5
44.4% · Extra high
FACTS Benchmark Suite
Counts toward:Weights: Research ×0.6Factuality across grounding, parametric, search and multimodal.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.1 Flash-Lite
40.6% · High
Humanity's Last Exam
Counts toward:Weights: Research ×0.5Accuracy on ~2,500 expert-written questions at the frontier of human knowledge (no tools unless noted).
Higher is better · 79 models tested · % solved · higher is better · 79 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
61.4% · Max
ProgramBench (fully resolved)
Counts toward:Weights: Coding ×0.5Share of 200 programs (from jq to SQLite/FFmpeg) an agent re-implements from only the binary and docs so that all hidden behavioral tests pass.
Higher is better · 41 models tested · % solved · higher is better · 41 models scored · the test's websiteofficial page ↗
Claude Sonnet 5.5
79.7% · Max
LMArena Text - Multi-Turn
Counts toward:Weights: Writing ×0.5Crowd-voted, style-controlled Elo on multi-turn conversations (iterative drafting/revision).
Higher is better · 28 models tested · Elo rating · higher is better · 28 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1552 · High
LMArena Text - Hard Prompts
Counts toward:Weights: Research ×0.5Crowd-voted, style-controlled Elo on prompts classified as hard (specific, complex, problem-solving).
Higher is better · 27 models tested · Elo rating · higher is better · 27 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1550 · High
LMArena Text (overall)
Counts toward:Weights: Writing ×0.5Crowd-voted Elo (style-controlled) from blind head-to-head chats across all text prompts.
Higher is better · 26 models tested · Elo rating · higher is better · 26 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1525 · High
NL2Repo-Bench
Counts toward:Weights: Coding ×0.5Generate an entire working repository from a natural-language specification, scored by the repo's test suite.
Higher is better · 19 models tested · % solved · higher is better · 19 models scored · the test's websiteofficial page ↗
DeepSeek V4.1 Flash
64.0% · Max
Terminal-Bench 3.0
Counts toward:Weights: Coding ×0.5Share of 74 frontier-difficulty terminal tasks (v0.1 rolling set, formerly Frontier-Bench) an agent completes; superseded by 4.0.
Higher is better · 18 models tested · % solved · higher is better · 18 models scored · the test's websiteofficial page ↗
Claude Opus 5
42.7% · Max
AA Analyst Agent
Counts toward:Weights: Research ×0.5Artificial Analysis agentic data-analyst tasks (pass@1).
Higher is better · 14 models tested · % solved · higher is better · 14 models scored · the test's websiteofficial page ↗
Gemini 3.7 Flash
60.0% · High
GDP.pdf
Counts toward:Weights: Research ×0.5Surge AI benchmark of professional questions about complex PDFs.
Higher is better · 13 models tested · % solved · higher is better · 13 models scored · the test's websiteofficial page ↗
Gemini 3.8 Flash
35.0% · Default
OpenAI MRCR v2 (8-needle)
Counts toward:Weights: Research ×0.5Multi-round coreference over long context; context range in settingLabel.
Higher is better · 13 models tested · % solved · higher is better · 13 models scored · the test's websiteofficial page ↗
GPT-6 Astra
100.0% · Default
Harvey's Legal Agent Benchmark (Vals)
Counts toward:Weights: Research ×0.5Complex legal research and drafting workflows; all-pass rate (Vals AI).
Higher is better · 6 models tested · % solved · higher is better · 6 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
19.6% · Max
Vals Finance Agent v2
Counts toward:Weights: Research ×0.5Multi-step financial analyst research tasks (Vals AI).
Higher is better · 6 models tested · % solved · higher is better · 6 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
65.4% · Max
OfficeQA Pro
Counts toward:Weights: Research ×0.5Questions over office documents (Pro split).
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
69.0% · Max
DeepSearchQA
Counts toward:Weights: Research ×0.5900 agentic browsing questions with list answers; F1.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
90.3% · Max
APEX-SWE
Counts toward:Weights: Coding ×0.5Mercor's professional software-engineering task benchmark.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Grok 4.6
56.4% · High
FrontierCode (Diamond)
Counts toward:Weights: Coding ×0.5Hardest FrontierCode subset (early version reported at Claude Fable 5 launch).
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Claude Fable 5
29.3% · Extra high
DeepSWE 1.0
Counts toward:Weights: Coding ×0.5Earlier version of Datacurve's long-horizon software engineering benchmark.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Grok 4.5
62.0% · Default
Writing codeCoding
Share of scientific-computing coding subproblems solved correctly (as run by Artificial Analysis).
Higher is better · 59 models tested · % solved · higher is better · 59 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
66.9% · Max
SWE-bench Verified
Too easy nowSaturatedShare of 500 human-validated real GitHub Python issues a model fixes so the repo tests pass.
Higher is better · 48 models tested · % solved · higher is better · 48 models scored · the test's websiteofficial page ↗
Claude Opus 5
97.0% · Default
LiveCodeBench
Too easy nowSaturatedPass@1 on recent competitive-programming problems from LeetCode, AtCoder and Codeforces (Vals implementation).
Higher is better · 34 models tested · % solved · higher is better · 34 models scored · the test's websiteofficial page ↗
Qwen3.8 Flash
91.9% · Default
Share of 300 real GitHub issues across 9 programming languages a model fixes (bash-only mini-SWE-agent runs).
Higher is better · 31 models tested · % solved · higher is better · 31 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
93.9% · Max
IOI (Vals)
Too easy nowSaturatedScore on International Olympiad in Informatics 2024-2026 problems solved agentically.
Higher is better · 25 models tested · % solved · higher is better · 25 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
100.0% · High
Share of contamination-free real-world binaries an agent can fully reverse-engineer.
Higher is better · 16 models tested · % solved · higher is better · 16 models scored · the test's websiteofficial page ↗
GPT-6 Astra
56.9% · Max
SWE-bench Verified (bash-only, mini-SWE-agent)
Too easy nowSaturatedSWE-bench Verified run by the SWE-bench team with the minimal bash-only mini-SWE-agent so models are compared on equal scaffolding.
Higher is better · 13 models tested · % solved · higher is better · 13 models scored · the test's websiteofficial page ↗
Claude Opus 4.5
76.8% · High
SWE-Bench Pro V2 (full)
Too easy nowSaturatedRefreshed 642-task SWE-Bench Pro public split with a locked, network-isolated protocol and re-grading on a pristine image.
Higher is better · 10 models tested · % solved · higher is better · 10 models scored · the test's websiteofficial page ↗
Claude Opus 5
99.4% · Extra high
Competitive programming rating on Codeforces problems (no tools).
Higher is better · 8 models tested · Elo rating · higher is better · 8 models scored · the test's websiteofficial page ↗
DeepSeek V4.1 Flash
3471 · Max
SWE-bench issues that include visual context such as screenshots and design mockups.
Higher is better · 8 models tested · % solved · higher is better · 8 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
61.4% · Max
Earlier version of the AA Coding Agent Index (different scale from v1.4).
Higher is better · 4 models tested · Score · higher is better · 4 models scored · the test's websiteofficial page ↗
GPT-5.6 Sol
80 · Default
CursorBench as reported in June 2026 system cards (version not stated); not comparable to later versions.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Claude Fable 5
72.9% · Max
ProgramBench reported as the share of tasks 'almost solved' (Almost@1); a much stricter metric than the default ProgramBench score.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
DeepSeek V4.1 Flash
20.3% · Max
Very long-horizon software engineering tasks (resolution rate, pass@1).
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
GLM-5.3
42.5% · Max
Moonshot in-house realistic coding-agent benchmark across 10+ languages (vendor-internal).
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Kimi K3
72.9% · Max
Anthropic-internal 478-problem subset of SWE-bench Pro used for cost/effort studies; not comparable to the public leaderboard.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
95.3% · High
Z.ai in-house coding-agent benchmark in realistic local dev environments, end-to-end completion rate (vendor-internal).
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
GLM-5.3
34.5% · Max
OpenAI internal long-horizon coding eval with ~20h median human completion time.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
GPT-5.5
73.1% · Extra high
Original FrontierCode v1 release (pre-v1.1), as reported in the Claude Sonnet 5 system card.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Claude Sonnet 5
38.8% · Max
Write fast GPU kernels; score is achieved TFLOPs relative to hardware peak (MiniMax setup).
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
MiniMax M3
28.8% · Default
Qwen in-house software-engineering benchmark run in Claude Code (vendor-internal, not comparable across vendors).
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Qwen3.8 Max (0803)
80.7% · Default
Security vulnerability remediation; pass@1.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 4 Argon
68.0% · Max
Competitive coding problems from Codeforces, ICPC and IOI, scored as Elo.
Higher is better · 1 model tested · Elo rating · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.1 Pro (Preview)
2887 · High
SWE-bench-style issue resolution across many programming languages.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
MiniMax M2.7
52.7% · Default
Make real repositories faster without breaking tests (performance-optimisation SWE tasks).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
MiniMax M3
34.8% · Default
MiniMax in-house app-building benchmark used at the M2.7 launch (vendor-internal).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
MiniMax M2.7
55.6% · Default
MiniMax in-house benchmark: build web/Android/iOS apps from scratch, verified by an agent-as-verifier (vendor-internal).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
MiniMax M3
50.1% · Default
Working on its ownAgentic
Terminal-Bench 2.1
Too easy nowSaturatedShare of 89 containerized command-line tasks an agent completes; a fixed revision of Terminal-Bench 2.0, now archived and near ceiling.
Higher is better · 68 models tested · % solved · higher is better · 68 models scored · the test's websiteofficial page ↗
GPT-5.6 Sol
91.9% · Default
Agent task success with and without reusable skill files, averaged; measures how well agents use packaged skills.
Higher is better · 30 models tested · % solved · higher is better · 30 models scored · the test's websiteofficial page ↗
Qwen3.8 Max (0803)
70.2% · Default
Long-horizon, economically valuable computer tasks spanning 55 sub-industries.
Higher is better · 26 models tested · % solved · higher is better · 26 models scored · the test's websiteofficial page ↗
GPT-6 Astra
59.3% · Max
Computer-use tasks in real desktop environments (361 tasks).
Higher is better · 21 models tested · % solved · higher is better · 21 models scored · the test's websiteofficial page ↗
Qwen3.8 Max (0803)
86.1% · Default
Binary success rate of computer-use agents on long-horizon real desktop tasks.
Higher is better · 18 models tested · % solved · higher is better · 18 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
77.9% · Max
70 scientific research workflows (data analysis, simulation, theorem proving) solved by an agent CLI and graded by hidden tests.
Higher is better · 10 models tested · % solved · higher is better · 10 models scored · the test's websiteofficial page ↗
GPT-6 Astra
68.1% · Max
Money balance after an agent runs a simulated vending-machine business for a year; tests long-horizon coherence.
Higher is better · 10 models tested · US dollars · higher is better · 10 models scored · the test's websiteofficial page ↗
GPT-6 Astra
$15,515 · Default
Find and trigger real vulnerabilities in open-source projects from source code (pass@1).
Higher is better · 9 models tested · % solved · higher is better · 9 models scored · the test's websiteofficial page ↗
DeepSeek V4.1 Flash
88.1% · Max
Length of software tasks (in human-expert minutes) an agent completes with 50% reliability.
Higher is better · 9 models tested · Time horizon · higher is better · 9 models scored · the test's websiteofficial page ↗
Claude Mythos Preview
17.4 h · Default
Internet-free subset of OSWorld 2.0 (v2026.08.08 release), partial reward, as reported by OpenAI.
Higher is better · 7 models tested · % solved · higher is better · 7 models scored · the test's websiteofficial page ↗
GPT-6 Astra
73.5% · Max
Kaggle-style machine-learning engineering competitions; average position score.
Higher is better · 5 models tested · % solved · higher is better · 5 models scored · the test's websiteofficial page ↗
Gemini 3.6 Flash
63.9% · Default
30-task subset of MLS-Bench: invent generalisable ML methods within 5 hours.
Higher is better · 5 models tested · % solved · higher is better · 5 models scored · the test's websiteofficial page ↗
Kimi K3
48.3% · Max
OSWorld 2.x as reported for the Claude 5.5 family (partial score).
Higher is better · 5 models tested · % solved · higher is better · 5 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
81.8% · Max
ML engineering: post-train base models within a 10-hour single-H100 budget.
Higher is better · 5 models tested · % solved · higher is better · 5 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
45.3% · Max
Long-horizon professional tasks (Mercor).
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Grok 4.6
57.5% · High
65 professional tasks across 35 white-collar occupations; mean rubric score.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
64.9% · Max
Vals AI financial-analysis agent benchmark (version noted per row).
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
GPT-5.5
60.0% · Extra high
OSWorld 2.0 tasks counted only if fully passed.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
41.7% · Max
Strict pass rate on the same tasks.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
48.7% · Max
Perplexity wide-search benchmark (500 tasks, structured fact collection with citations), soft F1.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
72.3% · Max
Meta-internal composite of instruction following in agentic tasks.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
57.8% · Max
Develop working exploits for real vulnerabilities; average capability-coverage score.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
GLM-5.3
54.4% · Max
Artificial Analysis' Elo rating of models on economically valuable knowledge-work tasks; vendor did not state which GDPval-AA version.
Higher is better · 2 models tested · Elo rating · higher is better · 2 models scored · the test's websiteofficial page ↗
Grok 4.7
1695 · Extra high
OSWorld 2.0 with strict binary task completion.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
32.0% · Max
Reproduce ML research papers end to end, graded with expert rubrics.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Qwen3.8 Max (0803)
93.0% · Default
Personal-assistant / OpenClaw-style agent tasks (pass^3 unless noted).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
MiniMax M3
74.5% · Default
General AI assistant agentic tasks (v2).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Muse Glimmer 30B
43.3% · High
Kaggle-style ML engineering competitions (lite subset), medal rate.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
MiniMax M2.7
66.6% · Default
Agent tasks in the OpenClaw-style personal agent harness.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Muse Glimmer 30B
47.6% · High
ReasoningReasoning
HLE with search/code tools (vendors blocklist HLE-discussing sources).
Higher is better · 32 models tested · % solved · higher is better · 32 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
67.7% · Max
Accuracy on trick multiple-choice questions about everyday spatial, temporal and social reasoning where ordinary humans do well.
Higher is better · 30 models tested · % solved · higher is better · 30 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
88.4% · Default
Vals.ai composite of real-world task benchmarks (finance, legal, coding, etc.).
Higher is better · 24 models tested · % solved · higher is better · 24 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
68.9% · High
Epoch AI composite capability score stitched across many benchmarks onto one scale.
Higher is better · 22 models tested · Score · higher is better · 22 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
167.3 · Default
ARC-AGI-2
Too easy nowSaturatedShare of novel abstract grid puzzles solved on the semi-private set; now near ceiling for frontier models.
Higher is better · 16 models tested · % solved · higher is better · 16 models scored · the test's websiteofficial page ↗
GPT-6 Astra
95.0% · Max
Score on interactive, game-like ARC-AGI-3 environments that require exploration and learning; vendor "provider adapter" harnesses score far higher than bare models.
Higher is better · 11 models tested · % solved · higher is better · 11 models scored · the test's websiteofficial page ↗
GPT-6 Astra
99.9% · High
ARC-AGI-1
Too easy nowSaturatedAbstract visual reasoning puzzles (v1).
Higher is better · 7 models tested · % solved · higher is better · 7 models scored · the test's websiteofficial page ↗
Claude Fable 5
98.5% · Max
Biology real-world research tasks, macro-average of 11 sub-tasks.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
88.8% · Max
Bioinformatics research workflows with tools; human-difficult split.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Gemini 3.8 Flash
56.5% · Default
Bioinformatics research workflows with tools; human-solvable split.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Gemini 3.8 Flash
88.8% · Default
1,811-item verified/revised subset of Humanity's Last Exam.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Gemini 3.8 Flash
54.9% · Default
Out-of-distribution precise instruction following.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Muse Glimmer 30B
77.0% · High
Harder successor to BIG-Bench Hard reasoning tasks.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Gemma 4 31B IT
74.4% · Default
Constrained text generation / instruction following.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Mistral Medium 3.5
95.8% · High
Electrical engineering tasks (xAI-reported).
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Grok 4.7
64.0% · Extra high
Condensed matter theory problems.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.1 Pro (Preview)
50.5% · Max
Research-level science problems.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Muse Spark
38.0% · Max
International Chemistry Olympiad 2025 theory problems.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.1 Pro (Preview)
82.8% · Max
International Physics Olympiad 2025 theory problems.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.1 Pro (Preview)
87.7% · Max
KnowledgeKnowledge
GPQA Diamond
Too easy nowSaturatedAccuracy on 198 graduate-level, Google-proof science multiple-choice questions.
Higher is better · 70 models tested · % solved · higher is better · 70 models scored · the test's websiteofficial page ↗
GPT-6 Astra
96.3% · Extra high
Share of AA-Omniscience questions answered correctly.
Higher is better · 32 models tested · % solved · higher is better · 32 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
67.2% · Max
LegalBench (Vals)
Too easy nowSaturatedOpen-source legal reasoning tasks (issue spotting, rules, interpretation, rhetoric).
Higher is better · 29 models tested · % solved · higher is better · 29 models scored · the test's websiteofficial page ↗
Claude Fable 5
88.6% · Max
Share of document summaries that contain claims not supported by the source document, judged by Vectara HHEM-2.3.
Lower is better · 28 models tested · % solved · lower is better · 28 models scored · the test's websiteofficial page ↗
Gemma 4 26B A4B IT
5.2% · Default
GDPval-AA version 2 (Elo; not comparable to v1 or v2.1).
Higher is better · 26 models tested · Elo rating · higher is better · 26 models scored · the test's websiteofficial page ↗
Claude Fable 5.1
1853 · Max
Swedish-language rank score across sentiment, NER, linguistic acceptability, reading comprehension, summarisation (SweDN), knowledge (Skolprov) and common-sense tasks; lower is better.
Lower is better · 25 models tested · Score · lower is better · 25 models scored · the test's websiteofficial page ↗
GPT-6 Astra
1.2 · Default
Vals-written tax questions graded on correctness and stepwise reasoning.
Higher is better · 25 models tested · % solved · higher is better · 25 models scored · the test's websiteofficial page ↗
Muse Spark 1.2
80.4% · Extra high
Clinician-oriented HealthBench split; Anthropic and OpenAI use different scoring (not comparable across vendors).
Higher is better · 14 models tested · % solved · higher is better · 14 models scored · the test's websiteofficial page ↗
Claude Sonnet 5.5
69.2% · Max
Accuracy on a curated, harder and cleaner subset of Humanity's Last Exam.
Higher is better · 12 models tested · % solved · higher is better · 12 models scored · the test's websiteofficial page ↗
GPT-6 Astra
60.6% · Default
Artificial Analysis agentic GDPval variant, Elo from blind pairwise comparisons.
Higher is better · 9 models tested · Elo rating · higher is better · 9 models scored · the test's websiteofficial page ↗
Claude Fable 5
1932 · Max
MMLU-Pro
Too easy nowSaturatedHarder 10-option version of MMLU.
Higher is better · 9 models tested · % solved · higher is better · 9 models scored · the test's websiteofficial page ↗
Qwen3.7 Max
89.6% · Default
MMMLU
Too easy nowSaturatedMultilingual MMLU.
Higher is better · 7 models tested · % solved · higher is better · 7 models scored · the test's websiteofficial page ↗
Gemini 3.1 Pro (Preview)
92.6% · High
Artificial Analysis long-horizon knowledge-work benchmark (multi-week projects), Elo; pre-v1.1.
Higher is better · 6 models tested · Elo rating · higher is better · 6 models scored · the test's websiteofficial page ↗
Claude Opus 5
1720 · Max
OpenAI's GDPval: expert-graded knowledge-work deliverables across 44 occupations (wins or ties vs experts).
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
GPT-5.5
84.9% · Extra high
MathsMath
FrontierMath (Tiers 1-3)
Too easy nowSaturatedAccuracy on hundreds of unpublished, expert-written research-level math problems (Epoch-run, v2).
Higher is better · 24 models tested · % solved · higher is better · 24 models scored · the test's websiteofficial page ↗
GPT-6.1 Sol
93.7% · Max
FrontierMath Tier 4
Too easy nowSaturatedAccuracy on the hardest FrontierMath tier of research-level problems (Epoch-run, v2).
Higher is better · 18 models tested · % solved · higher is better · 18 models scored · the test's websiteofficial page ↗
GPT-6.1 Sol
100.0% · Max
Share of problems from the Harvard-MIT Mathematics Tournament (Feb 2026) a model answers correctly.
Higher is better · 9 models tested · % solved · higher is better · 9 models scored · the test's websiteofficial page ↗
Qwen3.7 Max
97.1% · Default
Olympiad-level math problems with short verifiable answers.
Higher is better · 9 models tested · % solved · higher is better · 9 models scored · the test's websiteofficial page ↗
GLM-5.2
91.0% · Default
2026 American Invitational Mathematics Examination, no tools.
Higher is better · 8 models tested · % solved · higher is better · 8 models scored · the test's websiteofficial page ↗
GLM-5.2
99.2% · Default
MathArena's hardest set of recent competition problems (pass@1).
Higher is better · 7 models tested · % solved · higher is better · 7 models scored · the test's websiteofficial page ↗
DeepSeek V4.1 Flash
65.6% · Max
Epoch AI research-level math problems, tiers 1-3.
Higher is better · 5 models tested · % solved · higher is better · 5 models scored · the test's websiteofficial page ↗
GPT-5.6 Sol
89.0% · Default
MathArena final-answer problems extracted monthly from recent arXiv papers (release month noted per row).
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
91.2% · Max
ArXivMath with code execution.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
96.9% · Max
AIME 2025
Too easy nowSaturated2025 American Invitational Mathematics Examination.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Mistral Medium 3.5
86.3% · High
Harder-than-AIME competition math (avg@16).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Mistral Medium 3.5
66.9% · High
International Mathematical Olympiad 2025 problems.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.1 Pro (Preview)
81.5% · Max
Research-level mathematics (Surge leaderboard).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 4 Argon
76.0% · Max
Images and videoMultimodal
Multimodal college-level reasoning.
Higher is better · 17 models tested · % solved · higher is better · 17 models scored · the test's websiteofficial page ↗
Gemini 3.5 Flash
83.6% · Default
Information synthesis from complex scientific charts.
Higher is better · 10 models tested · % solved · higher is better · 10 models scored · the test's websiteofficial page ↗
Muse Spark
88.9% · Default
Surge AI chart-reading benchmark, no tools.
Higher is better · 5 models tested · % solved · higher is better · 5 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
71.6% · Max
CharXiv Reasoning with search/code tools.
Higher is better · 5 models tested · % solved · higher is better · 5 models scored · the test's websiteofficial page ↗
Gemini 3.6 Flash
89.4% · Default
Spatial reasoning from floor plans/blueprints.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Claude Fable 5
38.6% · Max
Extreme long video understanding (no tools).
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
91.7% · Max
MMMU-Pro with tool use.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
GPT-5.6 Sol
84.6% · Default
GUI grounding on high-resolution professional screenshots.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
GPT-6 Astra
92.7% · Extra high
Surge AI chart-reading benchmark with tools.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
89.0% · Max
Basic visual reasoning.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Muse Spark 1.1
76.3% · Default
Reconstruct 3D objects from multi-view renders by writing CAD code (geometric overlap).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
GPT-6 Astra
95.9% · Default
LVBench with Gemini's agentic video understanding.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.8 Flash
87.8% · Default
Document parsing benchmark (score as reported).
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Muse Glimmer 30B
75.8% · High
Knowledge acquisition from videos.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.1 Flash-Lite
84.8% · High
Long documentsLong context
Median output tokens per second on the first-party API, measured by Artificial Analysis (unit is tokens/sec).
Higher is better · 56 models tested · Tokens per second · higher is better · 56 models scored · the test's websiteofficial page ↗
Nemotron 3.5 Lightning 30B-A3B
298 tok/s · Default
MRCR v2 8-needle at 1M tokens.
Higher is better · 5 models tested · % solved · higher is better · 5 models scored · the test's websiteofficial page ↗
Gemini 3.6 Flash
54.0% · Default
MRCR v2 8-needle in the 256K-512K band.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
98.5% · Max
MRCR v2 8-needle in the 512K-1M band.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Muse Spark 1.3
98.1% · Max
Long-term memory benchmark at 128K.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Muse Glimmer 30B
65.1% · High
Graph traversal over 256k-1M contexts, BFS F1.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 4 Argon
84.2% · Max
Graph traversal over long contexts up to 128k, BFS F1.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 4 Argon
99.7% · Max
What people preferHuman preference
Crowd-voted, style-controlled Elo on all non-English prompts (no Swedish/Nordic category exists).
Higher is better · 27 models tested · Elo rating · higher is better · 27 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1512 · High
Crowd-voted, style-controlled Elo on business, management and financial-operations prompts.
Higher is better · 27 models tested · Elo rating · higher is better · 27 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1523 · High
Crowd-voted Elo (style-controlled) restricted to coding prompts in the text arena.
Higher is better · 23 models tested · Elo rating · higher is better · 23 models scored · the test's websiteofficial page ↗
Gemini 4 Argon
1559 · High
Crowd-voted Bradley-Terry Elo for generated websites, UI, games, slides and other design artifacts.
Higher is better · 21 models tested · Elo rating · higher is better · 21 models scored · the test's websiteofficial page ↗
Claude Opus 5.5
1399 · Default
Crowd-voted Elo for full-stack web apps built by models.
Higher is better · 19 models tested · Elo rating · higher is better · 19 models scored · the test's websiteofficial page ↗
GPT-6 Astra
1355 · Default
Crowd-voted Elo (raw) for search-grounded answers from models with web search/grounding enabled.
Higher is better · 18 models tested · Elo rating · higher is better · 18 models scored · the test's websiteofficial page ↗
GPT-5.6 Sol
1257 · Extra high
LLM-judged hard chat prompts.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Mistral Small 4
58.3% · High
Using toolsTool use
Share of realistic multi-step tasks solved by calling real MCP tool servers.
Higher is better · 36 models tested · % solved · higher is better · 36 models scored · the test's websiteofficial page ↗
Muse Spark 1.2
90.3% · Extra high
Zapier's benchmark of end-to-end business workflows across many connected apps (47 tools).
Higher is better · 33 models tested · % solved · higher is better · 33 models scored · the test's websiteofficial page ↗
DeepSeek V4.1 Flash
54.8% · Max
Multi-app tool-use tasks; Pass@1.
Higher is better · 18 models tested · % solved · higher is better · 18 models scored · the test's websiteofficial page ↗
Claude Opus 5
80.6% · Max
Verified subset of Toolathlon personal tool-use tasks.
Higher is better · 15 models tested · % solved · higher is better · 15 models scored · the test's websiteofficial page ↗
GLM-5.3-Flash
78.4% · Max
Success rate of agents handling simulated customer-service conversations with tools and policies.
Higher is better · 11 models tested · % solved · higher is better · 11 models scored · the test's websiteofficial page ↗
GLM-5.2
99.1% · Max
MCP tool-use tasks across Notion, GitHub, filesystem, Postgres and Playwright servers.
Higher is better · 4 models tested · % solved · higher is better · 4 models scored · the test's websiteofficial page ↗
Qwen3.7 Max
60.8% · Default
Human-verified edition of MCPMark.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Kimi K3
94.5% · Max
Customer-service agent simulation, telecom domain.
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Gemini 3.1 Pro (Preview)
99.3% · High
τ³-bench banking domain with agentic retrieval (Sierra).
Higher is better · 3 models tested · % solved · higher is better · 3 models scored · the test's websiteofficial page ↗
Muse Glimmer 30B
23.5% · High
Third-generation tau-bench customer-service agent simulations with policy following.
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Qwen3.6 Plus
70.7% · Default
τ³-bench airline domain (Sierra).
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Mistral Medium 3.5
72.0% · High
τ³-bench retail domain (Sierra).
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Mistral Medium 3.5
76.1% · High
τ³-bench telecom domain (Sierra).
Higher is better · 2 models tested · % solved · higher is better · 2 models scored · the test's websiteofficial page ↗
Mistral Medium 3.5
91.4% · High
τ²-bench retail domain.
Higher is better · 1 model tested · % solved · higher is better · 1 model scored · the test's websiteofficial page ↗
Gemini 3.1 Pro (Preview)
90.8% · High