Models · Anthropic · Out sinceReleased 17 Feb 2026
Claude Sonnet 4.6
Claude Sonnet 4.6 is made by Anthropic. We don't have enough test results yet to rank it. It's expensive to use.
API id claude-sonnet-4-6. Context/max output from OpenRouter listing. Anthropic recommends medium for agentic coding on Sonnet 4.6.
- —
- —
- —
- Expensive$3 / $15
- $6
- 1M
- 128K
- 30 (20 independent20 indep.)
- 17 Feb 2026
- Anthropic
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-sonnet-4.6.
Route it as anthropic/claude-sonnet-4.6 at $3 in / $15 out per 1M tokens, 1M context. Listed since 17 Feb 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_pverbosity
Thinking level
Reasoning effort
How long should Claude Sonnet 4.6Where Claude Sonnet 4.6think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
OSWorld 2.0
Best at Medium, worse abovePeaks at Medium
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AutomationBench | 5.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | |
| CursorBench (pre-3.2, legacy) | 49.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | Cursor agent | CursorBench version as of June 2026 (pre-3.2); effort per Cursor-reported best |
| Design Arena (all categories) | 1282 | Defaultclaude-sonnet-4-6 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.3 SE; 26314 battles; win rate 59.4% |
| Design Arena (fullstack) | 1211 | Defaultclaude-sonnet-4-6 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±8.6 SE; 2030 battles; win rate 64.3% |
| EQ-Bench Creative Writing v3 (Elo) | 1810 | Defaultclaude-sonnet-4-6 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 24; rubric score 16.50/20; slop 9.90; avg length 5876 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| EuroEval Swedish (generative) | 1.6 | Defaultclaude-sonnet-4-6 (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 4; ±0.05; lower is better; task scores (first metric): SweDN summarisation 36.50 ± 0.16, Skolprov 53.27 ± 4.13, Swedish facts 61.54 ± 2.33, ScaLA-sv 71.29 ± 1.37 |
| FrontierCode v1 | 15.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | Claude Code | |
| GDPval-AA v2 | 1395 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | |
| HealthBench Professional | 44.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | |
| Humanity's Last Exam | 34.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | no tools; regraded |
| Humanity's Last Exam (with tools) | 46.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | regraded |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 1.7 | DefaultClaude Sonnet 4.6 Thinking 16K | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 15/56; Thurstone comparison score (centered at 0); est. win chance 74%; 95% bootstrap 1.564 to 1.780 |
| LMArena Search Arena | 1221 | Defaultclaude-sonnet-4-6-search | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 7 (rank range 5-8); 95% CI 1216.4-1226.1; 134905 votes |
| LMArena Text - Coding category | 1529 | Defaultclaude-sonnet-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 24 (CI rank 7-42); 95% CI 1523-1534; 19706 votes |
| MCP Atlas | 69.5% | Defaultclaude-sonnet-4-6 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 13 (Scale rank accounts for CI); ±2.9; entry added 2025-11-06 |
| OSWorld 2.0 | 9.3% | Mediumclaude-sonnet-4-6_medium | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.33899999999999997; tool setting standard; step budget 500 |
| OSWorld 2.0 | 8.3% | Maxclaude-sonnet-4-6_max | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.41500000000000004; tool setting standard; step budget 500 |
| OSWorld-Verified | 78.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | updated methodology |
| ProgramBench (avg test pass rate) | 47.5% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 1.0%; avg cost $26.73/task |
| ProgramBench (fully resolved) | 0.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.5 stderr; strict fully-resolved rate |
| ProgramBench (fully resolved) | 0.0% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 1.0%; avg cost $26.73/task |
| SkillsBench | 49.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.64 stderr; $1.382/test |
| SWE Atlas - Codebase QnA | 31.2% | DefaultSonnet 4.6 (Claude Code) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 7 (Scale rank accounts for CI); ±5; entry added 2026-02-25 |
| SWE Atlas - Refactoring | 32.2% | DefaultSonnet-4.6 (Claude Code) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 9 (Scale rank accounts for CI); ±6.77; entry added 2026-05-06 |
| SWE Atlas - Test Writing | 31.8% | DefaultSonnet-4.6 (Claude Code) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 5 (Scale rank accounts for CI); ±6.24; entry added 2026-03-26 |
| SWE-Bench Pro (public, v1) | 58.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | — | |
| SWE-bench Verified | 75.2% | Defaultclaude-sonnet-4-6 | Independent testIndependentEpoch AI ↗ | 21 Feb 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 2.0pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| Terminal-Bench 2.1 | 67.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 30 Jun 2026 | mini-SWE-agent (GKE) | |
| Vals TaxEval v2 | 77.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id anthropic/claude-sonnet-4-6; rank 4/145; ±0.818 stderr; $0.127973/test |
| Vectara Hallucination Leaderboard (HHEM) | 10.6% | Defaultanthropic/claude-sonnet-4-6 | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 89.4 %; answer rate 99.9 %; avg summary 114.7 words; HHEM-2.3 judge; effort not stated (API default) |