Models · Anthropic · Out sinceReleased 28 May 2026
Claude Opus 4.8#17 for coding.#17 for coding, best at max effort.
Claude Opus 4.8 is made by Anthropic. Among the models we track it ranks #17 for coding, #21 for writing, #28 for research and analysis. It's expensive to use.
API id claude-opus-4-8. Default effort high; Anthropic recommends starting at xhigh for coding/agentic work. Fast mode $10/$50.
- 52.6 / 100 · #21
- 57.1 / 100 · #28
- 52.3 / 100 · #17
- Expensive$5 / $25
- $10
- 1M
- 128K
- 102 (67 independent67 indep.)
- 28 May 2026
- Anthropic
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-4.8.
Route it as anthropic/claude-opus-4.8 at $5 in / $25 out per 1M tokens, 1M context. Listed since 27 May 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolsverbosity
Thinking level
Reasoning effort
How long should Claude Opus 4.8Where Claude Opus 4.8think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
DeepSWE
Best at MaxBest at Max
FrontierCode
Best at MaxBest at Max
ProgramBench (fully resolved)
Best at MaxBest at Max
Terminal-Bench 3.0
Best at MaxBest at Max
Terminal-Bench 2.1
Best at MaxBest at Max
LLM Creative Story-Writing Benchmark (Lech Mazur)
Best at Extra highBest at Extra high
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase | 1346 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | |
| ARC-AGI-1 | 92.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | |
| ARC-AGI-2 | 72.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | |
| ARC-AGI-3 | 1.5% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | ARC Prize Foundation; read from launch-post benchmark table image (values printed in table) |
| Artificial Analysis Intelligence Index | 41.8 | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-4-8; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $4.08/task (AA marks this variant deprecated) |
| Artificial Analysis output speed | 59 tok/s | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 64.1s; list price $5/25 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 17.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| Blueprint-Bench 2 | 14.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 9 Jun 2026 | — | read from launch-post benchmark table image (values printed in table) |
| BrowseComp | 88.5% | Maxadaptive thinking, effort=max, multi-agent | Maker's own figureVendor-reportedAnthropic ↗ | 28 May 2026 | — | multi-agent |
| BrowseComp | 84.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| CursorBench (pre-3.2, legacy) | 63.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | Cursor agent | CursorBench version as of June 2026 (pre-3.2); effort per Cursor-reported best |
| DeepSWE | 54.4% | Extra highclaude-opus-4-8_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 80.5%; ±3.7; 4 runs; $8.01/task |
| DeepSWE | 59.0% | Maxclaude-opus-4-8_max | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 79.3%; ±1.8; 4 runs; $13.22/task |
| DeepSWE | 59.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| Design Arena (all categories) | 1259 | Defaultclaude-opus-4-8 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.2 SE; 29994 battles; win rate 52.8% |
| Design Arena (fullstack) | 1247 | Defaultclaude-opus-4-8 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±7.2 SE; 2698 battles; win rate 55.3% |
| Epoch Capabilities Index | 158.3 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 156.1-160.7; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1840 | Defaultclaude-opus-4-8 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 21; rubric score 16.66/20; slop 13.16; avg length 5842 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| Finance Agent | 53.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 May 2026 | — | Finance Agent v2 |
| Frontier-Bench v0.1 | 21.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| Frontier-Bench v0.1 | 18.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | mini-SWE-agent (GKE) | internal run |
| FrontierCode | 34.3% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 9 Jun 2026 | Claude Code | FrontierCode Main at Fable 5 launch (pass rate 37.3%) |
| FrontierCode | 46.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| FrontierCode | 46.5% | Defaultclaude-opus-4-8_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | Claude Code | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5 |
| FrontierCode (Diamond) | 13.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 9 Jun 2026 | Claude Code | read from launch-post benchmark table image (values printed in table) |
| FrontierMath (Tiers 1-3) | 80.0% | Maxclaude-opus-4-8_max | Independent testIndependentEpoch AI ↗ | 10 Jun 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 56.1% | Maxclaude-opus-4-8_max | Independent testIndependentEpoch AI ↗ | 10 Jun 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 7.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| GDP.pdf | 22.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 9 Jun 2026 | — | read from launch-post benchmark table image (values printed in table) |
| GDPval-AA (v1) | 1890 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 9 Jun 2026 | — | read from launch-post benchmark table image (values printed in table) |
| GDPval-AA v2 | 1593 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| GPQA Diamond | 92.4% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.984 stderr; $0.339/test |
| GPQA Diamond | 92.0% | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated) |
| GPQA Diamond | 93.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 May 2026 | — | |
| GSO | 47.1% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | 12 Jul 2026 | OpenHands | Opt@1; hack-controlled score 47.06; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable |
| HealthBench Professional | 57.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| Humanity's Last Exam | 48.7% | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated) |
| Humanity's Last Exam | 49.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| Humanity's Last Exam (with tools) | 57.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| LiveCodeBench | 87.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.948 stderr; $0.327/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.3 | HighClaude Opus 4.8 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 25/56; Thurstone comparison score (centered at 0); est. win chance 54%; 95% bootstrap 0.163 to 0.371; incomplete story set (see README coverage note) |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.8 | Extra highClaude Opus 4.8 (xhigh) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 18/56; Thurstone comparison score (centered at 0); est. win chance 62%; 95% bootstrap 0.666 to 0.888 |
| LMArena Search Arena | 1204 | Defaultclaude-opus-4-8 | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 12 (rank range 8-17); 95% CI 1197.9-1210.7; 70998 votes |
| LMArena Text - Coding category | 1533 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 16 (CI rank 7-31); 95% CI 1527-1539; 17275 votes |
| LMArena Text - Coding category | 1530 | Defaultclaude-opus-4-8 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 20 (CI rank 7-40); 95% CI 1524-1536; 17933 votes |
| LMArena Text - Creative Writing | 1469 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 16 (rank range 9-34); 95% CI 1462.8-1476.0; 12699 votes; style-controlled |
| LMArena Text - Expert | 1525 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 4-35); 95% CI 1517.4-1533.1; 7004 votes; style-controlled |
| LMArena Text - Expert | 1517 | Defaultclaude-opus-4-8 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 8-46); 95% CI 1508.8-1524.4; 7328 votes; style-controlled |
| LMArena Text - Hard Prompts | 1513 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 7-24); 95% CI 1508.3-1517.4; 42198 votes; style-controlled |
| LMArena Text - Instruction Following | 1490 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 6-21); 95% CI 1485.1-1495.8; 22488 votes; style-controlled |
| LMArena Text - Instruction Following | 1479 | Defaultclaude-opus-4-8 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 19 (rank range 13-39); 95% CI 1473.2-1483.9; 23069 votes; style-controlled |
| LMArena Text - Longer Query | 1502 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 7-26); 95% CI 1496.9-1507.2; 30035 votes; style-controlled |
| LMArena Text - Longer Query | 1500 | Defaultclaude-opus-4-8 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 7-28); 95% CI 1494.6-1504.9; 30720 votes; style-controlled |
| LMArena Text - Multi-Turn | 1498 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 10 (rank range 7-33); 95% CI 1491.0-1504.4; 11116 votes; style-controlled |
| LMArena Text - Multi-Turn | 1495 | Defaultclaude-opus-4-8 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 7-39); 95% CI 1487.8-1501.3; 11326 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1489 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 3-32); 95% CI 1482.9-1495.7; 12641 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1487 | Defaultclaude-opus-4-8 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 16 (rank range 6-36); 95% CI 1480.3-1493.1; 12693 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1495 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 16 (rank range 3-50); 95% CI 1486.4-1504.2; 5385 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1497 | Defaultclaude-opus-4-8 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 2-46); 95% CI 1488.0-1505.7; 5472 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1475 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 9-36); 95% CI 1468.7-1480.4; 17150 votes; style-controlled |
| LMArena Text (overall) | 1481 | Highclaude-opus-4-8-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 23 (CI rank 13-40); 95% CI 1477-1485; 63577 votes |
| MCP Atlas | 82.2% | Maxclaude-opus-4-8 (max) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±2.4; entry added 2026-05-28 |
| MCP Atlas | 82.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 May 2026 | — | |
| OSWorld 2.0 | 20.6% | Maxclaude-opus-4-8_max | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.5479999999999999; tool setting batched tool; step budget 500 |
| OSWorld 2.0 | 55.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| OSWorld-Verified | 83.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 9 Jun 2026 | — | read from launch-post benchmark table image (values printed in table) |
| ProgramBench (avg test pass rate) | 70.9% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 16.5%; avg cost $21.02/task |
| ProgramBench (fully resolved) | 0.0% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 16.5%; avg cost $21.02/task |
| ProgramBench (fully resolved) | 1.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.705 stderr; $31.267/test; strict fully-resolved rate |
| SciCode | 54.4% | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated) |
| ScreenSpot-Pro | 82.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 May 2026 | — | no tools |
| SimpleBench | 64.8% | DefaultClaude Opus 4.8 | Independent testIndependentSimpleBench ↗ | 29 May 2026 | — | AVG@5, temp 0.7; rank 23rd; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 53.0% | Maxclaude-opus-4-8_max | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SkillsBench | 59.2% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.5 stderr; $4.767/test |
| SWE Atlas - Codebase QnA | 57.3% | Extra highOpus 4.8 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±4.93; entry added 2026-06-08 |
| SWE Atlas - Refactoring | 46.7% | DefaultOpus 4.8 (Claude Code) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±6.75; entry added 2026-06-08 |
| SWE Atlas - Test Writing | 49.6% | Extra highOpus 4.8 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±5.93; entry added 2026-06-08 |
| SWE-bench Multilingual | 84.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | |
| SWE-bench Multimodal | 38.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | |
| SWE-Bench Pro (public, v1) | 69.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | avg of 5 trials |
| SWE-bench Verified | 88.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.423 stderr; $1.923/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| SWE-bench Verified | 88.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 May 2026 | — | |
| SWE-bench Verified | 85.8% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | Claude Code | ±1.563 stderr; $0.672/test; Vals variant run through Claude Code harness; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| tau2-bench | 94.4% | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 82.7% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 9 Jun 2026 | mini-SWE-agent (GKE) | |
| Terminal-Bench 2.1 | 74.6% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 May 2026 | Terminus-2 (Harbor, Daytona) | 89 tasks x 5 attempts |
| Terminal-Bench 2.1 | 84.6% | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 71.9% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.649 stderr; $2.410/test |
| Terminal-Bench 2.1 | 78.9% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 3.0 | 15.0% | Highshipped default effort (high) | Maker's own figureVendor-reportedAnthropic ↗ | 1 Oct 2026 | Claude Managed Agents | 74 tasks, two runs per model; $63 per solved task |
| Terminal-Bench 3.0 | 21.1% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 5.2B, run cost $5.2k |
| Terminal-Bench 4.0 | 23.6% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.56 95% CI; 330 trials; total run cost $6481; model release 2026-05-28 |
| Terminal-Bench 4.0 | 23.2% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.336 stderr; $17.141/test |
| Terminal-Bench 4.0 | 21.7% | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated) |
| Toolathlon | 79.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | internal harness, Pass@1 |
| Vals Code Migration | 47.3% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.176 stderr; $30.514/test |
| Vals CorpFin v2 | 66.7% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id anthropic/claude-opus-4-8; rank 18/134; ±0.928 stderr; $0.855534/test |
| Vals Index | 55.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.009 stderr; $13.136/test |
| Vals Legal Research Bench | 43.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-opus-4-8; rank 19/72; ±3.448 stderr; $2.820942/test |
| Vals Public Benefits Bench v1.1 | 68.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-opus-4-8; rank 11/45; ±1.212 stderr; $1.889795/test |
| Vals TaxEval v2 | 75.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id anthropic/claude-opus-4-8; rank 14/145; ±0.838 stderr; $0.159139/test |
| Vals Vibe Code Bench | 82.7% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±3.081 stderr; $26.880/test |
| Vals Vibe Code Bench | 77.5% | Default | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | Claude Code | ±3.737 stderr; $6.856/test; Vals variant run through Claude Code harness |