Skip to content
Bencher

Models · Anthropic · Out sinceReleased 4 Feb 2026

Claude Opus 4.6#5 for writing.#5 for writing, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Claude Opus 4.6 is made by Anthropic. Among the models we track it ranks #5 for writing, #21 for research and analysis. It's expensive to use.

Auto-created from OpenRouter catalog; verify details.

Writing & creativity
71.2 / 100 · #5
Research & analysis
62.9 / 100 · #21
Coding
—
Price
Expensive$5 / $25
Price per 1M (blended)Blended / 1M
$10
MemoryContext
1M
Longest answerMax output
—
Test resultsResults
55 (55 independent55 indep.)
Out sinceReleased
4 Feb 2026
Made byVendor
Anthropic
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-4.6.

Route it as anthropic/claude-opus-4.6 at $5 in / $25 out per 1M tokens, 1M context. Listed since 4 Feb 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_pverbosity

Thinking level

Reasoning effort

How long should Claude Opus 4.6Where Claude Opus 4.6think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Humanity's Last Exam

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Design Arena (all categories)1295Defaultclaude-opus-4-6Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.3 SE; 27400 battles; win rate 61.2%
Design Arena (fullstack)1218Defaultclaude-opus-4-6Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±7.6 SE; 2513 battles; win rate 59.8%
EQ-Bench Creative Writing v3 (Elo)1809Defaultclaude-opus-4-6Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 25; rubric score 16.53/20; slop 12.12; avg length 6055 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
GSO41.2%Highreasoning_effort=highIndependent testIndependentGSO ↗27 Apr 2026OpenHandsOpt@1; hack-controlled score 37.25; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
GSO33.3%DefaultIndependent testIndependentGSO ↗11 Feb 2026OpenHandsOpt@1; hack-controlled score 30.39; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
Humanity's Last Exam19.0%No reasoningclaude-opus-4-6 (Non-Thinking)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 16 (Scale rank accounts for CI); ±1.54; entry added 2026-02-17
Humanity's Last Exam34.4%Maxclaude-opus-4-6-thinking-maxIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 6 (Scale rank accounts for CI); ±1.86; entry added 2026-02-17
LegalBench (Vals)85.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-opus-4-6-thinking; rank 20/149; ±0.368 stderr; $0.004043/test
LLM Creative Story-Writing Benchmark (Lech Mazur)1.1DefaultClaude Opus 4.6 Thinking 16KIndependent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 17/56; Thurstone comparison score (centered at 0); est. win chance 67%; 95% bootstrap 0.974 to 1.333
LMArena Search Arena1253Defaultclaude-opus-4-6-searchIndependent testIndependentLMArena ↗24 Aug 2026—rank 2 (rank range 1-2); 95% CI 1248.4-1258.4; 134699 votes
LMArena Text - Coding category1551Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 3 (CI rank 1-12); 95% CI 1546-1557; 20144 votes
LMArena Text - Coding category1547Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 5 (CI rank 1-13); 95% CI 1542-1553; 22791 votes
LMArena Text - Creative Writing1501Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 1-6); 95% CI 1494.8-1507.7; 13974 votes; style-controlled
LMArena Text - Creative Writing1479Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 6-22); 95% CI 1472.3-1484.9; 14136 votes; style-controlled
LMArena Text - Expert1547Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-13); 95% CI 1539.0-1555.3; 7012 votes; style-controlled
LMArena Text - Expert1533Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 1-23); 95% CI 1525.7-1540.9; 8390 votes; style-controlled
LMArena Text - Hard Prompts1533Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 2-7); 95% CI 1529.2-1537.6; 49235 votes; style-controlled
LMArena Text - Hard Prompts1527Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 2-10); 95% CI 1522.4-1530.5; 53159 votes; style-controlled
LMArena Text - Instruction Following1514Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 1-4); 95% CI 1508.5-1518.9; 24809 votes; style-controlled
LMArena Text - Instruction Following1499Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 3-15); 95% CI 1494.3-1504.3; 27380 votes; style-controlled
LMArena Text - Longer Query1524Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 2-7); 95% CI 1519.0-1528.9; 32696 votes; style-controlled
LMArena Text - Longer Query1517Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 2-9); 95% CI 1512.1-1521.7; 35553 votes; style-controlled
LMArena Text - Multi-Turn1519Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 2-8); 95% CI 1512.1-1524.8; 13183 votes; style-controlled
LMArena Text - Multi-Turn1512Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 2-13); 95% CI 1505.7-1517.9; 14496 votes; style-controlled
LMArena Text - Non-English1491Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 2-13); 95% CI 1487.2-1495.6; 43047 votes; style-controlled
LMArena Text - Non-English1485Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 2-20); 95% CI 1480.7-1488.9; 45139 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1501Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 1-15); 95% CI 1495.2-1507.1; 15275 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1500Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 1-15); 95% CI 1494.5-1506.1; 16232 votes; style-controlled
LMArena Text - Occupational: Legal & Government1511Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 1-27); 95% CI 1502.7-1519.5; 6312 votes; style-controlled
LMArena Text - Occupational: Legal & Government1506Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 1-32); 95% CI 1498.1-1514.6; 6572 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1500Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 2-8); 95% CI 1494.8-1505.8; 19810 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1490Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 4-16); 95% CI 1484.3-1495.2; 20192 votes; style-controlled
LMArena Text (overall)1506Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 2 (CI rank 2-7); 95% CI 1502-1509; 77193 votes
LMArena Text (overall)1497Defaultclaude-opus-4-6Independent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 7 (CI rank 3-15); 95% CI 1494-1501; 81769 votes
MCP Atlas76.8%Maxclaude-opus-4-6 (max)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±2.7; entry added 2025-12-18
METR 50% time horizon12.0 hDefaultIndependent testIndependentMETR ↗8 May 2026METR react agent (Inspect)50% time horizon, METR-Horizon-v1.1; 95% CI 317-3634 min; 80% horizon 69.9 min; release 2026-02-05; raw data https://metr.org/assets/benchmark_results_1_1.yaml
PRBench Finance (Scale)53.3%No reasoningclaude-opus-4-6 (Non-Thinking)Independent testIndependentScale AI (SEAL) ↗17 Feb 2026—Scale rank 4; ±0.1802 CI
PRBench Legal (Scale)52.3%No reasoningclaude-opus-4-6 (Non-Thinking)Independent testIndependentScale AI (SEAL) ↗17 Feb 2026—Scale rank 3; ±0.6555 CI
ProgramBench (avg test pass rate)52.1%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 2.5%; avg cost $11.38/task
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 2.5%; avg cost $11.38/task
SimpleBench67.6%DefaultClaude Opus 4.6Independent testIndependentSimpleBench ↗17 Feb 2026—AVG@5, temp 0.7; rank 20th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified47.0%Maxclaude-opus-4-6_maxIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SWE Atlas - Codebase QnA33.3%DefaultOpus 4.6 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 7 (Scale rank accounts for CI); ±5; entry added 2026-02-25
SWE Atlas - Codebase QnA30.0%DefaultOpus 4.6 (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 11 (Scale rank accounts for CI); ±4.9; entry added 2026-02-25
SWE Atlas - Refactoring35.6%DefaultOpus-4.6 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 9 (Scale rank accounts for CI); ±6.82; entry added 2026-05-06
SWE Atlas - Test Writing36.7%DefaultOpus-4.6 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 3 (Scale rank accounts for CI); ±6.63; entry added 2026-03-26
SWE Atlas - Test Writing36.1%DefaultOpus-4.6 (Mini-SWE)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 3 (Scale rank accounts for CI); ±6.02; entry added 2026-03-26
SWE-bench Multilingual72.0%DefaultIndependent testIndependentSWE-bench ↗13 Feb 2026mini-SWE-agent 2.0.0a0Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26
SWE-Bench Pro (private/commercial set)47.1%Defaultclaude-opus-4-6 (thinking)*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 1 (Scale rank accounts for CI); ±6.07; entry added 2026-04-08; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)51.9%Defaultclaude-opus-4-6 (thinking)*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±3.61; entry added 2026-04-08; * = Scale footnote (see page)
SWE-bench Verified78.7%Defaultclaude-opus-4-6Independent testIndependentEpoch AI ↗18 Feb 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)
SWE-bench Verified (bash-only, mini-SWE-agent)75.6%DefaultIndependent testIndependentSWE-bench ↗17 Feb 2026mini-SWE-agent 2.0.0Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26
Vals CorpFin v267.0%Maxeffort=maxIndependent testIndependentVals.ai ↗12 Aug 2026—vals id anthropic/claude-opus-4-6-thinking; rank 14/134; ±0.926 stderr; $0.487461/test
Vals TaxEval v276.0%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id anthropic/claude-opus-4-6-thinking; rank 8/145; ±0.834 stderr; $0.071368/test
Vectara Hallucination Leaderboard (HHEM)12.2%Defaultanthropic/claude-opus-4-6Independent testIndependentVectara ↗22 Sep 2026—factual consistency 87.8 %; answer rate 99.8 %; avg summary 137.6 words; HHEM-2.3 judge; effort not stated (API default)