Models · Anthropic · Out sinceReleased 30 Jun 2026
Claude Sonnet 5#27 for coding.#27 for coding, best at max effort.
Claude Sonnet 5 is made by Anthropic. Among the models we track it ranks #27 for coding, #29 for research and analysis, #31 for writing. It's mid-priced to use.
API id claude-sonnet-5. Introductory $2/$10 pricing made permanent Aug 10, 2026. Default effort high. New tokenizer (~1.0-1.35x tokens vs Sonnet 4.6).
- 32.6 / 100 · #31
- 55.8 / 100 · #29
- 36.0 / 100 · #27
- Mid-priced$2 / $10
- $4
- 1M
- 128K
- 92 (50 independent50 indep.)
- 30 Jun 2026
- Anthropic
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-sonnet-5.
Route it as anthropic/claude-sonnet-5 at $2 in / $10 out per 1M tokens, 1M context. Listed since 30 Jun 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstool_choicetoolsverbosity
Thinking level
Reasoning effort
How long should Claude Sonnet 5Where Claude Sonnet 5think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
FrontierCode
Best at Extra high, worse abovePeaks at Extra high
Terminal-Bench 2.1
Best at MaxBest at Max
SciCode
Best at MaxBest at Max
Artificial Analysis Intelligence Index
Best at MaxBest at Max
AA-Briefcase v1.1
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 923 | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $0.82 |
| AA-Briefcase v1.1 | 1056 | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $1.73 |
| AA-Briefcase v1.1 | 1177 | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $3.76 |
| AA-Briefcase v1.1 | 1274 | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $7.56 |
| AA-Briefcase v1.1 | 1359 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $14.43 |
| Artificial Analysis Intelligence Index | 34.4 | Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-sonnet-5-xhigh; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $2.87/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 38.2 | MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-sonnet-5; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $5.09/task (AA marks this variant deprecated) |
| Artificial Analysis output speed | 63 tok/s | Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 20.2s; list price $2/10 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 79 tok/s | MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 212.4s; list price $2/10 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 10.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| BrowseComp | 86.6% | Maxadaptive thinking, effort=max, multi-agent | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | multi-agent configuration |
| BrowseComp | 84.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | single agent; 10M token limit with compaction |
| Chartography (no tools) | 15.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | no tools |
| CursorBench | 24.1% | LowSonnet 5 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 61; $1.39/task; 23,772 tokens/task; 46 steps/task |
| CursorBench | 24.1% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $1.39 |
| CursorBench | 28.0% | MediumSonnet 5 Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 54; $2.31/task; 39,114 tokens/task; 65 steps/task |
| CursorBench | 28.0% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $2.31 |
| CursorBench | 30.8% | HighSonnet 5 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 50; $3.48/task; 61,146 tokens/task; 85 steps/task |
| CursorBench | 30.8% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $3.48 |
| CursorBench | 32.0% | Extra highSonnet 5 Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 47; $4.55/task; 83,373 tokens/task; 102 steps/task |
| CursorBench | 32.0% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $4.55 |
| CursorBench | 34.1% | MaxSonnet 5 Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 39; $7.17/task; 149,257 tokens/task; 140 steps/task |
| CursorBench | 34.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $7.17 |
| CursorBench (pre-3.2, legacy) | 61.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | Cursor agent | CursorBench version as of June 2026 (pre-3.2); effort per Cursor-reported best |
| Design Arena (all categories) | 1283 | Defaultclaude-sonnet-5 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.5 SE; 21694 battles; win rate 53.2% |
| Design Arena (fullstack) | 1247 | Defaultclaude-sonnet-5 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±6.3 SE; 3733 battles; win rate 53.7% |
| EQ-Bench Creative Writing v3 (Elo) | 1794 | Defaultclaude-sonnet-5 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 27; rubric score 16.47/20; slop 11.60; avg length 5753 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| Frontier-Bench v0.1 | 17.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | mini-SWE-agent (GKE) | internal run |
| FrontierCode | 28.7% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $2.39 |
| FrontierCode | 35.2% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $3.81 |
| FrontierCode | 39.4% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $6.1 |
| FrontierCode | 42.7% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $10.07 |
| FrontierCode | 42.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $17.12 |
| FrontierCode | 42.7% | Defaultclaude-sonnet-5_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | Claude Code | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5 |
| FrontierCode v1 | 38.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | Claude Code | |
| GDPval-AA v2 | 1618 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | Elo as of June 17, 2026 |
| GDPval-AA v2.1 | 1449 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations) |
| GPQA Diamond | 91.1% | MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated) |
| GSO | 37.3% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | 12 Jul 2026 | OpenHands | Opt@1; hack-controlled score 36.27; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable |
| HealthBench Professional | 57.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | |
| Humanity's Last Exam | 39.0% | Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-xhigh) (AA marks this variant deprecated) |
| Humanity's Last Exam | 41.3% | MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated) |
| Humanity's Last Exam | 43.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | no tools |
| Humanity's Last Exam (with tools) | 54.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| LegalBench (Vals) | 83.9% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-sonnet-5; rank 44/149; ±0.458 stderr; $0.005759/test |
| LMArena Search Arena | 1194 | Defaultclaude-sonnet-5-search | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 17 (rank range 12-18); 95% CI 1187.2-1201.0; 40230 votes |
| LMArena Text - Creative Writing | 1436 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 61 (rank range 36-84); 95% CI 1428.4-1443.1; 9284 votes; style-controlled |
| LMArena Text - Expert | 1512 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 23 (rank range 8-54); 95% CI 1502.9-1520.6; 5051 votes; style-controlled |
| LMArena Text - Hard Prompts | 1491 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 45 (rank range 28-60); 95% CI 1485.8-1495.6; 29825 votes; style-controlled |
| LMArena Text - Instruction Following | 1467 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 37 (rank range 20-58); 95% CI 1461.1-1472.7; 15961 votes; style-controlled |
| LMArena Text - Longer Query | 1482 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 38 (rank range 21-59); 95% CI 1476.3-1487.4; 21121 votes; style-controlled |
| LMArena Text - Multi-Turn | 1474 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 51 (rank range 18-78); 95% CI 1466.0-1481.1; 7677 votes; style-controlled |
| LMArena Text - Non-English | 1448 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 55 (rank range 42-77); 95% CI 1443.2-1453.2; 26108 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1465 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 48 (rank range 25-77); 95% CI 1457.7-1472.0; 8766 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1478 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 41 (rank range 13-89); 95% CI 1467.6-1487.9; 3876 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1452 | Highclaude-sonnet-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 46 (rank range 29-70); 95% CI 1445.8-1458.8; 12341 votes; style-controlled |
| OSWorld 2.1 (partial score) | 57.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | partial score |
| OSWorld-Verified | 81.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | updated OSWorld-Verified methodology (zoom-tool bug fix, 128K max tokens/turn) |
| ProgramBench (fully resolved) | 0.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0 stderr; $24.358/test; strict fully-resolved rate |
| ProgramBench (fully resolved) | 77.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| SciCode | 54.1% | Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-xhigh) (AA marks this variant deprecated) |
| SciCode | 54.3% | MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated) |
| SimpleBench | 60.6% | DefaultClaude Sonnet 5 | Independent testIndependentSimpleBench ↗ | 9 Jul 2026 | — | AVG@5, temp 0.7; rank 34th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SWE-bench Multilingual | 78.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| SWE-bench Multimodal | 28.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| SWE-bench Pro (Anthropic internal subset) | 77.4% | Highadaptive thinking, effort=high (default) | Maker's own figureVendor-reportedAnthropic ↗ | 1 Oct 2026 | — | Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.84 per solved task |
| SWE-Bench Pro (public, v1) | 63.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| SWE-Bench Pro V2 (full) | 93.2% | Extra highSonnet 5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 8 (Scale rank accounts for CI); ±1.71; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-Bench Pro V2 (hard) | 88.2% | Extra highSonnet 5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 4 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-bench Verified | 85.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | |
| SWE-rebench | 56.8% | HighSonnet 5 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | 1 Oct 2026 | SWE-rebench standard scaffold | time window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±0.94; pass@5 74.8%; $1.43/problem |
| Terminal-Bench 2.1 | 80.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | mini-SWE-agent (GKE) | 89 tasks x 5 attempts |
| Terminal-Bench 2.1 | 80.5% | MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 74.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±2.085 stderr; $0.535/test |
| Terminal-Bench 2.1 | 74.6% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 3.0 | 14.6% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 17.9B, run cost $6.9k |
| Terminal-Bench 4.0 | 3.2% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $3.78 |
| Terminal-Bench 4.0 | 4.4% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $4.83 |
| Terminal-Bench 4.0 | 4.5% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $8.2 |
| Terminal-Bench 4.0 | 7.1% | Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-xhigh) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 7.0% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $9.95 |
| Terminal-Bench 4.0 | 14.1% | MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 12.4% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.06 95% CI; 330 trials; total run cost $9604; model release 2026-06-30 |
| Terminal-Bench 4.0 | 10.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $11.62 |
| Toolathlon | 74.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | internal harness, Pass@1 |
| Vals Code Migration | 44.4% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.246 stderr; $35.310/test |
| Vals CorpFin v2 | 67.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id anthropic/claude-sonnet-5; rank 15/134; ±0.926 stderr; $0.516119/test |
| Vals Index | 51.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.092 stderr; $13.721/test |
| Vals Legal Research Bench | 41.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-sonnet-5; rank 20/72; ±3.428 stderr; $2.719446/test |
| Vals Public Benefits Bench v1.1 | 66.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-sonnet-5; rank 17/45; ±1.232 stderr; $1.290695/test |
| Vals TaxEval v2 | 75.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id anthropic/claude-sonnet-5; rank 15/145; ±0.836 stderr; $0.301619/test |
| Vals Vibe Code Bench | 81.3% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±3.046 stderr; $25.385/test |