Skip to content
Bencher

Models · Anthropic · Out sinceReleased 24 Nov 2025

Claude Opus 4.5#23 for writing.#23 for writing, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Claude Opus 4.5 is made by Anthropic. Among the models we track it ranks #23 for writing. It's expensive to use.

Auto-created from OpenRouter catalog; verify details.

Writing & creativity
50.3 / 100 · #23
Research & analysis
—
Coding
—
Price
Expensive$5 / $25
Price per 1M (blended)Blended / 1M
$10
MemoryContext
200K
Longest answerMax output
—
Test resultsResults
15 (15 independent15 indep.)
Out sinceReleased
24 Nov 2025
Made byVendor
Anthropic
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-4.5.

Route it as anthropic/claude-opus-4.5 at $5 in / $25 out per 1M tokens, 200K context. Listed since 24 Nov 2025.

include_reasoningmax_completion_tokensmax_tokensreasoningresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_kverbosity

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Humanity's Last Exam25.2%Defaultclaude-opus-4-5-20251101-thinkingIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 10 (Scale rank accounts for CI); ±1.7; entry added 2025-11-26
LMArena Search Arena1180Defaultclaude-opus-4-5-searchIndependent testIndependentLMArena ↗24 Aug 2026—rank 19 (rank range 18-22); 95% CI 1174.0-1185.3; 61573 votes
LMArena Text - Coding category1531Highclaude-opus-4-5-20251101-high-32kIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 19 (CI rank 7-42); 95% CI 1523-1538; 7790 votes
LMArena Text - Coding category1523Defaultclaude-opus-4-5-20251101Independent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 27 (CI rank 8-53); 95% CI 1518-1528; 17972 votes
LMArena Text - Creative Writing1469Highclaude-opus-4-5-20251101-high-32kIndependent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 7-35); 95% CI 1460.8-1477.4; 5677 votes; style-controlled
LMArena Text - Instruction Following1483Highclaude-opus-4-5-20251101-high-32kIndependent testIndependentLMArena ↗30 Sep 2026—rank 17 (rank range 8-35); 95% CI 1476.5-1489.8; 9867 votes; style-controlled
LMArena Text - Longer Query1495Highclaude-opus-4-5-20251101-high-32kIndependent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 9-37); 95% CI 1488.5-1501.7; 9830 votes; style-controlled
MCP Atlas69.8%Highclaude-opus-4-5 (high)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 13 (Scale rank accounts for CI); ±2.9; entry added 2025-09-10
METR 50% time horizon4.9 hDefaultIndependent testIndependentMETR ↗8 May 2026METR react agent (Inspect)50% time horizon, METR-Horizon-v1.1; 95% CI 162-624 min; 80% horizon 49.4 min; release 2025-11-24; raw data https://metr.org/assets/benchmark_results_1_1.yaml
PRBench Finance (Scale)46.2%Defaultclaude-opus-4-5-20251101-thinkingIndependent testIndependentScale AI (SEAL) ↗26 Nov 2025—Scale rank 13; ±0.27 CI
PRBench Legal (Scale)44.2%Defaultclaude-opus-4-5-20251101-thinkingIndependent testIndependentScale AI (SEAL) ↗26 Nov 2025—Scale rank 18; ±0.34 CI
SWE-bench Multilingual70.7%DefaultIndependent testIndependentSWE-bench ↗13 Feb 2026mini-SWE-agent 2.0.0a0Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26
SWE-Bench Pro (public, v1)45.9%Defaultclaude-opus-4-5-20251101Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 5 (Scale rank accounts for CI); ±3.6; entry added 2025-12-11
SWE-bench Verified (bash-only, mini-SWE-agent)76.8%HighIndependent testIndependentSWE-bench ↗17 Feb 2026mini-SWE-agent 2.0.0Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26
Vectara Hallucination Leaderboard (HHEM)10.9%Defaultanthropic/claude-opus-4-5-20251101Independent testIndependentVectara ↗22 Sep 2026—factual consistency 89.1 %; answer rate 98.7 %; avg summary 114.5 words; HHEM-2.3 judge; effort not stated (API default)