Models · Anthropic · Out sinceReleased 4 Feb 2026
Claude Opus 4.6#5 for writing.#5 for writing, best at high effort.
Claude Opus 4.6 is made by Anthropic. Among the models we track it ranks #5 for writing, #21 for research and analysis. It's expensive to use.
Auto-created from OpenRouter catalog; verify details.
- 71.2 / 100 · #5
- 62.9 / 100 · #21
- —
- Expensive$5 / $25
- $10
- 1M
- —
- 55 (55 independent55 indep.)
- 4 Feb 2026
- Anthropic
- —
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-4.6.
Route it as anthropic/claude-opus-4.6 at $5 in / $25 out per 1M tokens, 1M context. Listed since 4 Feb 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_pverbosity
Thinking level
Reasoning effort
How long should Claude Opus 4.6Where Claude Opus 4.6think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Humanity's Last Exam
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Design Arena (all categories) | 1295 | Defaultclaude-opus-4-6 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.3 SE; 27400 battles; win rate 61.2% |
| Design Arena (fullstack) | 1218 | Defaultclaude-opus-4-6 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±7.6 SE; 2513 battles; win rate 59.8% |
| EQ-Bench Creative Writing v3 (Elo) | 1809 | Defaultclaude-opus-4-6 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 25; rubric score 16.53/20; slop 12.12; avg length 6055 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| GSO | 41.2% | Highreasoning_effort=high | Independent testIndependentGSO ↗ | 27 Apr 2026 | OpenHands | Opt@1; hack-controlled score 37.25; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable |
| GSO | 33.3% | Default | Independent testIndependentGSO ↗ | 11 Feb 2026 | OpenHands | Opt@1; hack-controlled score 30.39; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable |
| Humanity's Last Exam | 19.0% | No reasoningclaude-opus-4-6 (Non-Thinking) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 16 (Scale rank accounts for CI); ±1.54; entry added 2026-02-17 |
| Humanity's Last Exam | 34.4% | Maxclaude-opus-4-6-thinking-max | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 6 (Scale rank accounts for CI); ±1.86; entry added 2026-02-17 |
| LegalBench (Vals) | 85.3% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-opus-4-6-thinking; rank 20/149; ±0.368 stderr; $0.004043/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 1.1 | DefaultClaude Opus 4.6 Thinking 16K | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 17/56; Thurstone comparison score (centered at 0); est. win chance 67%; 95% bootstrap 0.974 to 1.333 |
| LMArena Search Arena | 1253 | Defaultclaude-opus-4-6-search | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 2 (rank range 1-2); 95% CI 1248.4-1258.4; 134699 votes |
| LMArena Text - Coding category | 1551 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 3 (CI rank 1-12); 95% CI 1546-1557; 20144 votes |
| LMArena Text - Coding category | 1547 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 5 (CI rank 1-13); 95% CI 1542-1553; 22791 votes |
| LMArena Text - Creative Writing | 1501 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 4 (rank range 1-6); 95% CI 1494.8-1507.7; 13974 votes; style-controlled |
| LMArena Text - Creative Writing | 1479 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 6-22); 95% CI 1472.3-1484.9; 14136 votes; style-controlled |
| LMArena Text - Expert | 1547 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-13); 95% CI 1539.0-1555.3; 7012 votes; style-controlled |
| LMArena Text - Expert | 1533 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 9 (rank range 1-23); 95% CI 1525.7-1540.9; 8390 votes; style-controlled |
| LMArena Text - Hard Prompts | 1533 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 2-7); 95% CI 1529.2-1537.6; 49235 votes; style-controlled |
| LMArena Text - Hard Prompts | 1527 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 5 (rank range 2-10); 95% CI 1522.4-1530.5; 53159 votes; style-controlled |
| LMArena Text - Instruction Following | 1514 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 3 (rank range 1-4); 95% CI 1508.5-1518.9; 24809 votes; style-controlled |
| LMArena Text - Instruction Following | 1499 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 6 (rank range 3-15); 95% CI 1494.3-1504.3; 27380 votes; style-controlled |
| LMArena Text - Longer Query | 1524 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 3 (rank range 2-7); 95% CI 1519.0-1528.9; 32696 votes; style-controlled |
| LMArena Text - Longer Query | 1517 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 5 (rank range 2-9); 95% CI 1512.1-1521.7; 35553 votes; style-controlled |
| LMArena Text - Multi-Turn | 1519 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 3 (rank range 2-8); 95% CI 1512.1-1524.8; 13183 votes; style-controlled |
| LMArena Text - Multi-Turn | 1512 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 2-13); 95% CI 1505.7-1517.9; 14496 votes; style-controlled |
| LMArena Text - Non-English | 1491 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 5 (rank range 2-13); 95% CI 1487.2-1495.6; 43047 votes; style-controlled |
| LMArena Text - Non-English | 1485 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 10 (rank range 2-20); 95% CI 1480.7-1488.9; 45139 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1501 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 6 (rank range 1-15); 95% CI 1495.2-1507.1; 15275 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1500 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 1-15); 95% CI 1494.5-1506.1; 16232 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1511 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 6 (rank range 1-27); 95% CI 1502.7-1519.5; 6312 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1506 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 1-32); 95% CI 1498.1-1514.6; 6572 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1500 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 4 (rank range 2-8); 95% CI 1494.8-1505.8; 19810 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1490 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 8 (rank range 4-16); 95% CI 1484.3-1495.2; 20192 votes; style-controlled |
| LMArena Text (overall) | 1506 | Highclaude-opus-4-6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 2 (CI rank 2-7); 95% CI 1502-1509; 77193 votes |
| LMArena Text (overall) | 1497 | Defaultclaude-opus-4-6 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 7 (CI rank 3-15); 95% CI 1494-1501; 81769 votes |
| MCP Atlas | 76.8% | Maxclaude-opus-4-6 (max) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±2.7; entry added 2025-12-18 |
| METR 50% time horizon | 12.0 h | Default | Independent testIndependentMETR ↗ | 8 May 2026 | METR react agent (Inspect) | 50% time horizon, METR-Horizon-v1.1; 95% CI 317-3634 min; 80% horizon 69.9 min; release 2026-02-05; raw data https://metr.org/assets/benchmark_results_1_1.yaml |
| PRBench Finance (Scale) | 53.3% | No reasoningclaude-opus-4-6 (Non-Thinking) | Independent testIndependentScale AI (SEAL) ↗ | 17 Feb 2026 | — | Scale rank 4; ±0.1802 CI |
| PRBench Legal (Scale) | 52.3% | No reasoningclaude-opus-4-6 (Non-Thinking) | Independent testIndependentScale AI (SEAL) ↗ | 17 Feb 2026 | — | Scale rank 3; ±0.6555 CI |
| ProgramBench (avg test pass rate) | 52.1% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 2.5%; avg cost $11.38/task |
| ProgramBench (fully resolved) | 0.0% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 2.5%; avg cost $11.38/task |
| SimpleBench | 67.6% | DefaultClaude Opus 4.6 | Independent testIndependentSimpleBench ↗ | 17 Feb 2026 | — | AVG@5, temp 0.7; rank 20th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 47.0% | Maxclaude-opus-4-6_max | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SWE Atlas - Codebase QnA | 33.3% | DefaultOpus 4.6 (Claude Code) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 7 (Scale rank accounts for CI); ±5; entry added 2026-02-25 |
| SWE Atlas - Codebase QnA | 30.0% | DefaultOpus 4.6 (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 11 (Scale rank accounts for CI); ±4.9; entry added 2026-02-25 |
| SWE Atlas - Refactoring | 35.6% | DefaultOpus-4.6 (Claude Code) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 9 (Scale rank accounts for CI); ±6.82; entry added 2026-05-06 |
| SWE Atlas - Test Writing | 36.7% | DefaultOpus-4.6 (Claude Code) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 3 (Scale rank accounts for CI); ±6.63; entry added 2026-03-26 |
| SWE Atlas - Test Writing | 36.1% | DefaultOpus-4.6 (Mini-SWE) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 3 (Scale rank accounts for CI); ±6.02; entry added 2026-03-26 |
| SWE-bench Multilingual | 72.0% | Default | Independent testIndependentSWE-bench ↗ | 13 Feb 2026 | mini-SWE-agent 2.0.0a0 | Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26 |
| SWE-Bench Pro (private/commercial set) | 47.1% | Defaultclaude-opus-4-6 (thinking)* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 1 (Scale rank accounts for CI); ±6.07; entry added 2026-04-08; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 51.9% | Defaultclaude-opus-4-6 (thinking)* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±3.61; entry added 2026-04-08; * = Scale footnote (see page) |
| SWE-bench Verified | 78.7% | Defaultclaude-opus-4-6 | Independent testIndependentEpoch AI ↗ | 18 Feb 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| SWE-bench Verified (bash-only, mini-SWE-agent) | 75.6% | Default | Independent testIndependentSWE-bench ↗ | 17 Feb 2026 | mini-SWE-agent 2.0.0 | Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26 |
| Vals CorpFin v2 | 67.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id anthropic/claude-opus-4-6-thinking; rank 14/134; ±0.926 stderr; $0.487461/test |
| Vals TaxEval v2 | 76.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id anthropic/claude-opus-4-6-thinking; rank 8/145; ±0.834 stderr; $0.071368/test |
| Vectara Hallucination Leaderboard (HHEM) | 12.2% | Defaultanthropic/claude-opus-4-6 | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 87.8 %; answer rate 99.8 %; avg summary 137.6 words; HHEM-2.3 judge; effort not stated (API default) |