Skip to content
Bencher

Models · Anthropic · Out sinceReleased 17 Feb 2026

Claude Sonnet 4.6

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Claude Sonnet 4.6 is made by Anthropic. We don't have enough test results yet to rank it. It's expensive to use.

API id claude-sonnet-4-6. Context/max output from OpenRouter listing. Anthropic recommends medium for agentic coding on Sonnet 4.6.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Expensive$3 / $15
Price per 1M (blended)Blended / 1M
$6
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
30 (20 independent20 indep.)
Out sinceReleased
17 Feb 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-sonnet-4.6.

Route it as anthropic/claude-sonnet-4.6 at $3 in / $15 out per 1M tokens, 1M context. Listed since 17 Feb 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_pverbosity

Thinking level

Reasoning effort

How long should Claude Sonnet 4.6Where Claude Sonnet 4.6think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

OSWorld 2.0

Best at Medium, worse abovePeaks at Medium

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AutomationBench5.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—
CursorBench (pre-3.2, legacy)49.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026Cursor agentCursorBench version as of June 2026 (pre-3.2); effort per Cursor-reported best
Design Arena (all categories)1282Defaultclaude-sonnet-4-6Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.3 SE; 26314 battles; win rate 59.4%
Design Arena (fullstack)1211Defaultclaude-sonnet-4-6Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±8.6 SE; 2030 battles; win rate 64.3%
EQ-Bench Creative Writing v3 (Elo)1810Defaultclaude-sonnet-4-6Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 24; rubric score 16.50/20; slop 9.90; avg length 5876 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
EuroEval Swedish (generative)1.6Defaultclaude-sonnet-4-6 (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 4; ±0.05; lower is better; task scores (first metric): SweDN summarisation 36.50 ± 0.16, Skolprov 53.27 ± 4.13, Swedish facts 61.54 ± 2.33, ScaLA-sv 71.29 ± 1.37
FrontierCode v115.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026Claude Code
GDPval-AA v21395Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—
HealthBench Professional44.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—
Humanity's Last Exam34.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—no tools; regraded
Humanity's Last Exam (with tools)46.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—regraded
LLM Creative Story-Writing Benchmark (Lech Mazur)1.7DefaultClaude Sonnet 4.6 Thinking 16KIndependent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 15/56; Thurstone comparison score (centered at 0); est. win chance 74%; 95% bootstrap 1.564 to 1.780
LMArena Search Arena1221Defaultclaude-sonnet-4-6-searchIndependent testIndependentLMArena ↗24 Aug 2026—rank 7 (rank range 5-8); 95% CI 1216.4-1226.1; 134905 votes
LMArena Text - Coding category1529Defaultclaude-sonnet-4-6Independent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 24 (CI rank 7-42); 95% CI 1523-1534; 19706 votes
MCP Atlas69.5%Defaultclaude-sonnet-4-6Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 13 (Scale rank accounts for CI); ±2.9; entry added 2025-11-06
OSWorld 2.09.3%Mediumclaude-sonnet-4-6_mediumIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.33899999999999997; tool setting standard; step budget 500
OSWorld 2.08.3%Maxclaude-sonnet-4-6_maxIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.41500000000000004; tool setting standard; step budget 500
OSWorld-Verified78.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—updated methodology
ProgramBench (avg test pass rate)47.5%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 1.0%; avg cost $26.73/task
ProgramBench (fully resolved)0.5%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.5 stderr; strict fully-resolved rate
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 1.0%; avg cost $26.73/task
SkillsBench49.0%Maxeffort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.64 stderr; $1.382/test
SWE Atlas - Codebase QnA31.2%DefaultSonnet 4.6 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 7 (Scale rank accounts for CI); ±5; entry added 2026-02-25
SWE Atlas - Refactoring32.2%DefaultSonnet-4.6 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 9 (Scale rank accounts for CI); ±6.77; entry added 2026-05-06
SWE Atlas - Test Writing31.8%DefaultSonnet-4.6 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 5 (Scale rank accounts for CI); ±6.24; entry added 2026-03-26
SWE-Bench Pro (public, v1)58.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—
SWE-bench Verified75.2%Defaultclaude-sonnet-4-6Independent testIndependentEpoch AI ↗21 Feb 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 2.0pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)
Terminal-Bench 2.167.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026mini-SWE-agent (GKE)
Vals TaxEval v277.1%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id anthropic/claude-sonnet-4-6; rank 4/145; ±0.818 stderr; $0.127973/test
Vectara Hallucination Leaderboard (HHEM)10.6%Defaultanthropic/claude-sonnet-4-6Independent testIndependentVectara ↗22 Sep 2026—factual consistency 89.4 %; answer rate 99.9 %; avg summary 114.7 words; HHEM-2.3 judge; effort not stated (API default)