Skip to content
Bencher

Which AI should you use?

Current SOTA pick, by use case

Claude Opus 5.5.

Available to everyone. On the expensive side.

anthropic/claude-opus-5.5 · $4 / $20 per 1M · on OpenRouter

Why?Rationale & comparison

Why this model

Rationale

Why —and how it compares.

The best writer you can use today

Best for writing & creativity

Claude Opus 5.5at High thinkingat high effort

In blind tests where thousands of people compare two anonymous texts, people prefer Claude Opus 5.5's writing over every other model you can use today. It's also the strongest at long, well-built stories and texts.

Opus 5.5 (high) is the top generally available model in LMArena's style-controlled Creative Writing (1515) and Writing & Literature (1519) categories, trailing only the trusted-tester-only Gemini 4 Argon (1521.8 / 1523, within the error bars). It is statistically tied for #1 on Lech Mazur's story-writing benchmark (3.77 vs Fable 5.1's 3.80). Available on OpenRouter; $4/$20 per 1M tokens.

How to use it. Choose Claude Opus 5.5 and set thinking to High. For writing, more thinking doesn't make the text better — it just takes longer.

Config. Model anthropic/claude-opus-5.5, reasoning effort high via OpenRouter.

Google's new Gemini 4 Argon is just ahead in the same tests, but it's only open to a small group of testers so far. And as always: read through anything you send out.

No evidence that xhigh/max improves writing; LMArena lists Opus 5.5 at high. Swedish: EuroEval's Swedish board doesn't cover Claude 5.x yet (GPT-6 Astra leads it at 1.19, lower is better).

Questions people ask

FAQ

Score for writing, out of 100

Writing & creativity composite (0–100)

  1. Claude Opus 5.5High thinkingHighOur pickPick

    87

    The best writer you can use today

    Best for writing & creativity

    LLM Creative Story-Writing Benchmark (Lech Mazur) 3.8EQ-Bench Creative Writing v3 (Elo) 2050LMArena Text - Creative Writing 1515

    Expensive$4 / $20 per 1MAvailable on OpenRouterOpenRouter

  2. Gemini 4 ArgonHigh thinkingHighPreview — not available yetPreview · no OpenRouter

    98

    Slightly ahead in people's votes, but you can't use it yet — only a small test group has access.

    #1 on LMArena Creative Writing (1521.8) and Writing & Literature (1523), but trusted-tester preview only and not on OpenRouter.

    LMArena Text - Creative Writing 1522

    Mid-priced · about half the price of Claude Opus 5.5$2 / $10 per 1MNot on OpenRouterNot on OpenRouter

  3. Claude Fable 5.1High thinkingHigh

    75

    Just as good at long creative pieces, but two and a half times the price.

    #1 on Lech Mazur story-writing (3.80) and EQ-Bench Creative Writing v3 (2162), but $10/$50 and lower on LMArena (1481.6).

    LLM Creative Story-Writing Benchmark (Lech Mazur) 3.8EQ-Bench Creative Writing v3 (Elo) 2162LMArena Text - Creative Writing 1482

    Expensive · about 3× the price of Claude Opus 5.5$10 / $50 per 1MAvailable on OpenRouterOpenRouter

  4. GPT-6 AstraStandard thinkingDefault

    58

    An AI judge rates its stories highest, but real people prefer Claude's writing.

    Highest EQ-Bench Creative Writing v3 score (2173.3, LLM-judged) but well behind on human-preference LMArena Creative Writing (1448.3).

    LLM Creative Story-Writing Benchmark (Lech Mazur) 3.5EQ-Bench Creative Writing v3 (Elo) 2173LMArena Text - Creative Writing 1448

    Expensive · about 3× the price of Claude Opus 5.5$10 / $50 per 1MAvailable on OpenRouterOpenRouter

  5. Claude Fable 5High thinkingHigh

    78

    Also strong for writing, at about 3× the price of Claude Opus 5.5.

    Writing & creativity composite 78.2 across 8 benchmarks (92% of weight).

    LLM Creative Story-Writing Benchmark (Lech Mazur) 2.9EQ-Bench Creative Writing v3 (Elo) 1943LMArena Text - Creative Writing 1503

    Expensive · about 3× the price of Claude Opus 5.5$10 / $50 per 1MAvailable on OpenRouterOpenRouter

100 = best of all the models we track on these tests, 0 = the weakest.

Weighted mean of min–max normalised benchmarks, best result per model at any setting. ↓ = lower is better.

Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Effort maps to each vendor's reasoning control (reasoning.effort, thinking budget). Higher isn't always better: some models overthink.

Checked 1 Oct 2026.Updated 1 Oct 2026 · Initial dataset: 121 models, ~4,400 cited scores across writing, analysis and coding from vendor posts and independent leaderboards. What changedChangelog

In the apps you already havePer-platform configuration

Using Claude, Gemini or ChatGPT at work?What to select in each appHere's what to pick.by use case.

Your workplace chat app decides which models you can choose. Here's the best option in each one, and how to set it. Tap a choice to see why.Exact model-picker label and thinking setting per platform and use case, with each app's reasoning control and data terms. Click a cell for the rationale.

Claude (Team plan)
How to change the thinking levelReasoning control, plans & data terms

Click the model name next to the send button, open "Effort" and pick a level (Low, Medium, High, Extra high, Max). Higher effort gives more thorough answers but is slower and uses up your limit faster. On Opus 5.5, Sonnet 5.5 and Fable 5.1 thinking is always on, so effort is the only dial.

Model menu next to the send button > "Effort" > Low / Medium / High / Extra high (xhigh) / Max. These are the same five levels as the API effort parameter (low/medium/high/xhigh/max); Extra high is only on Opus 4.7 and newer. Each model's recommended level is marked "Default" in the menu; the help article does not say which level that is per model (API defaults: Opus 5.5 = medium, Fable 5.1 = high, Sonnet 5.5 = high; that the app uses the same defaults is unconfirmed). Thinking is a separate "Thinking" (or "Extended") toggle under Effort, but it can't be turned off on Sonnet 5.5, Opus 5.5, Fable 5.1 or Opus 5 (adaptive thinking is always on). Older models without effort have only the "Extended" toggle.

Claude Team plan (Standard and Premium seats) on web, desktop and mobile. Opus/Sonnet/Haiku are included on both seat types; Fable models are included only on Premium seats (up to 50% of the weekly limit) and run on pay-as-you-go usage credits on Standard seats. Premium seats have 5x the usage of Standard seats.

Team (and Enterprise): 'No model training on your content by default' (claude.com/pricing). Memory is off by default for Team orgs (release notes 2026-08-25).

Checked 1 Oct 2026 · support.claude.com, support.claude.com, claude.com, support.claude.com, platform.claude.com, support.claude.com, anthropic.com

Writing

Opus 5.5Effort: High

Pick Opus 5.5 and set Effort to High. It is the best writer you can use today.

Opus 5.5 @ high: top generally available model on LMArena Creative Writing (1515) and Writing & Literature (1519). Effort menu maps 1:1 to API effort.

Research & analysis

Opus 5.5Effort: Max

Pick Opus 5.5 and set Effort to Max. Slower, but the analysis gets clearly better — and it makes fewer things up.

Opus 5.5 @ max: GDPval-AA v2.1 1846 vs 1692 at high; hallucination rate 58.6% vs 67.6% at high. Uses the weekly limit faster.

Coding

Opus 5.5Effort: High

Pick Opus 5.5 with Effort on High. Higher settings cost more and don't write better code.

Opus 5.5 @ high: CursorBench 56.0 vs 57.8 at max, Terminal-Bench 4.0 64.2 vs 64.8 at max, at ~1/3 the cost; FrontierCode drops at xhigh (51.4).

Gemini in Google Workspace
How to change the thinking levelReasoning control, plans & data terms

Click the model name inside the text box and pick Flash-Lite, Flash or Pro. Then click the model name again, open "Thinking level" and choose Standard (the default, faster) or Extended (thinks longer for hard problems). "Deep Think" only shows up with the AI Ultra add-on and requires Pro.

Model menu inside the composer: Flash-Lite / Flash / Pro, plus a "Thinking level" sub-menu with Standard (default), Extended, and Deep Think (AI Ultra only, Pro model only, takes minutes). Google does not document how Standard and Extended map to the API's thinking_level. For reference, the API's thinking_level is low/medium/high: gemini-3.1-pro-preview defaults to high and gemini-3.8-flash to medium. The Standard/Extended help text comes from the personal-account article. The Workspace edition table still uses the older 'Pro / Thinking / Fast' rows. Whether Workspace accounts see the exact same Thinking-level menu is unconfirmed.

The Gemini app on work Google accounts. Business Standard and Business Plus (and Enterprise Standard/Plus) get 'Pro access … with enterprise-grade security & privacy'. Business Starter gets 'Standard access'. Higher limits need the paid add-ons AI Expanded Access ('Expanded' badge) and AI Ultra Access (formerly Google AI Ultra for Business, 'Ultra' badge). Gemini in Docs/Gmail (side panel) has separate limits.

Workspace Business/Enterprise editions with Gemini as a core service: 'Your chats and uploaded files in Gemini Apps won't be reviewed by human reviewers or otherwise used to improve generative AI models' (Google Workspace Terms). This doesn't apply when Gemini is an 'additional service' (e.g. Business Base, Workspace Individual).

Checked 1 Oct 2026 · support.google.com, support.google.com, support.google.com, support.google.com, gemini.google, blog.google, deepmind.google, workspaceupdates.googleblog.com, ai.google.dev

Writing

ProThinking level: Standard

Pick Pro with the standard thinking level. Good, but a step behind Claude for writing until Google's newest model arrives.

Gemini 3.1 Pro: LMArena Creative Writing 1480 (vs 1515 Opus 5.5). Pro is capped at 25 prompts / 4 h on Business Standard/Plus.

Research & analysis

ProThinking level: Extended

Pick Pro and switch thinking to Extended for long documents and careful analysis.

Gemini 3.1 Pro: AA-LCR 82.0 (Opus 5.5: 84.7), hallucination rate 50.9%. Mapping of Standard/Extended to API thinking_level is undocumented.

Coding

ProThinking level: Extended

Pick Pro with Extended thinking — but for real coding work, Claude or ChatGPT's Codex are clearly stronger.

Gemini 3.1 Pro (vendor): SWE-bench Verified 80.6, SWE-bench Pro 54.2 — a generation behind current coding agents. Gemini 3.8 Flash (stronger at coding) isn't confirmed for Workspace.

ChatGPT (Business plan)
How to change the thinking levelReasoning control, plans & data terms

In a normal chat, use the model picker or Thinking slider: Instant (fast), then Medium, High and Extra High (more thinking), and Pro for the hardest tasks (limited monthly or weekly allowance). In ChatGPT Work there is a 'Power' slider from Faster to Smarter. Under 'Advanced' you can pick the model and its reasoning effort yourself.

Chat: Instant (GPT-5.6 Sol, fast) / Medium (GPT-5.6 Sol, 'standard reasoning') / High ('extended reasoning') / Extra High ('highest selectable reasoning effort') / Pro (GPT-5.6 Sol Pro or GPT-6 Pro = GPT-6 Astra). The labels match API reasoning_effort medium/high/xhigh, but OpenAI does not document an exact mapping. Business setting 'Higher intelligence' (Settings > General) lets Instant add reasoning automatically. Plus has Medium and High only (no Extra High or Pro). Work/Codex: a Power slider (Faster↔Smarter) with presets like Luna High, GPT-6 Sol Light/Medium, Astra Light/Medium/Extra High. 'Advanced' picks the model plus effort Light/Medium/High/Extra High/Max/Ultra. Light = API 'low'. Ultra uses subagents. Start with the default and raise it for hard tasks.

ChatGPT Business (Standard and Premium seats). Plus is noted where it differs. ChatGPT now has three areas: Chat (ordinary conversations), ChatGPT Work and Codex. Each has its own model picker and allowances.

Business: 'By default, we do not use data from … ChatGPT Business … for training or improving our models'. Plus is an individual plan and may be used for training unless the user turns off 'Improve the model for everyone' (Settings > Data controls).

Checked 1 Oct 2026 · help.openai.com, help.openai.com, learn.chatgpt.com, openai.com, help.openai.com

Writing

HighThinking slider: High

In a normal chat, choose High. Good everyday writing — though Claude still writes better.

GPT-5.6 Sol: LMArena Creative Writing 1468.9, Writing & Literature 1482 (measured at xhigh). No evidence that Extra High improves prose.

Research & analysis

Extra HighThinking slider: Extra High

Choose Extra High for analysis. Save Pro (limited messages) for the hardest questions.

GPT-5.6 Sol: GDPval-AA v2.1 1548 at xhigh vs 1480 at high (max 1588 isn't exposed in Chat). Pro = GPT-6 Astra / 5.6 Sol Pro, 15 msgs/month on Standard seats.

Coding

GPT-6.1 Sol (Codex / Work)Effort: Extra High

Use Codex (or ChatGPT Work) with GPT-6.1 Sol on Extra High. Don't go to Max — it gets worse.

GPT-6.1 Sol @ xhigh: DeepSWE 73.2 vs 69.6 at max; AA Coding Agent Index 62.9 (xhigh) vs 60.1 (max). Not selectable in plain Chat.

Other good choicesOther verdicts

Need something specific?Not just the headline picks.We've got a pick for that.Picks for every job.

Great code for half the price

Best coding value

Claude Sonnet 5.5

Anthropic · Maximum thinkingMax

Nearly as good as the top pick on coding tests and costs half as much per word. Choose Claude Sonnet 5.5 and set thinking to the maximum. It needs the extra thinking to do its best work.

Sonnet 5.5 at max tops Artificial Analysis' Coding Agent Index (68.4) and its Terminal-Bench 4.0 run (66.2) at $2/$10 per 1M tokens — half of Opus 5.5. It is very effort-sensitive, so the low list price only pays off at max.

Mid-priced$2 / $10 per 1MOn OpenRouterOpenRouter

Best for long, multi-step tasks

Best for long-running agents

Claude Fable 5.1

Anthropic · Maximum thinkingMax

When an AI has to work on its own for a long time — using a computer, browsing, or handling a whole task end-to-end — this one is the most reliable. Choose Claude Fable 5.1. It's expensive, so save it for big jobs where it works on its own.

Fable 5.1 leads computer-use and knowledge-work agent evals: OSWorld 2.0 77.9 and GDPval-AA v2 1853 (vendor-reported), MCP Atlas 87.2 (independent). Anthropic positions it for the hardest long-horizon work; it costs $10/$50 per 1M tokens.

Expensive$10 / $50 per 1MOn OpenRouterOpenRouter

Best for the hardest thinking problems

Best for hard reasoning

GPT-6 Astra

OpenAI · High thinkingHigh

It solves the toughest logic and science puzzles better than any other model. Choose GPT-6 Astra and set reasoning to High. Going higher doesn't help.

GPT-6 Astra tops ARC-AGI-2 (95.0) and ARC-AGI-3 (99.9 at high) on ARC Prize's verified leaderboard, plus GPQA Diamond 96.3. On ARC-AGI-3, high is its peak — xhigh and max score lower.

Expensive$10 / $50 per 1MOn OpenRouterOpenRouter

Fast and cheap, still very capable

Best fast & cheap

Gemini 3.8 Flash

Google (Gemini / DeepMind) · High thinkingHigh

For quick, everyday tasks it's close to the best models but costs a fraction and answers very fast. Choose Gemini 3.8 Flash with thinking set to High.

Gemini 3.8 Flash scores 73.8 on DeepSWE (independent) — within a point of the frontier — at $0.75/$3.75 per 1M tokens (intro price until 2026-12-31) and ~221 output tokens/s.

Mid-priced$0.75 / $3.75 per 1MOn OpenRouterOpenRouter

The best model you can run on your own servers

Best open model to run yourself

GLM-5.3

Z.ai (Zhipu) · Maximum thinkingMax

GLM-5.3 is the strongest freely downloadable model that still fits on a single AI server. Its licence lets you use it commercially, so your documents never have to leave your own or an EU data centre. Download GLM-5.3 from Z.ai and run it on one 8-GPU server, with thinking set to the maximum. If that's too big, the Flash version is half the size and even better at Swedish.

GLM-5.3 is #1 among open models on Artificial Analysis' Coding Agent Index (53.6) and Terminal-Bench 4.0 (41.9), with an AA Intelligence Index of 44.8 (max). At 744B total / 40B active it needs ~780 GB at 8-bit, so it fits one 8×H200 node. The licence is modified MIT (commercial use allowed; only API providers with more than $10B revenue need a security review).

See all open models →All open-weights models →
Mid-priced$1.40 / $4.40 per 1MOn OpenRouterOpenRouter

RankingsLeaderboard

Every model, rankedEvery model, rankedfrom best to worst.per use case.

Pick an area to see how all the models we track compare. Each model gets one score out of 100, based on the tests for that area.One composite per model (or per model × setting) for each use case, blended from that use case's weighted benchmarks. Switch to every setting to see how effort changes the ranking.

writing leaderboard
Best thinking levelMostly at
1Gemini 4 Argon97.7Google (Gemini / DeepMind)High thinkingHigh——152215521523Mid-priced$2 / $10—
2Claude Opus 5.586.6AnthropicHigh thinkingHigh3.82050151514971519Expensive$4 / $20
3Claude Fable 578.2AnthropicHigh thinkingHigh2.91943150315181509Expensive$10 / $50
4Claude Fable 5.174.9AnthropicMaximum thinkingMax3.82162148214871493Expensive$10 / $50
5Claude Opus 4.671.2AnthropicHigh thinkingHigh1.11809150115191500Expensive$5 / $25
6Claude Opus 571.1AnthropicHigh thinkingHigh3.821331472—1482Expensive$5 / $25
7Claude Opus 4.769.5AnthropicHigh thinkingHigh1.81914148915161497Expensive$5 / $25
8Kimi K361.5Moonshot AI (Kimi)Maximum thinkingMax2.52082146015001476Expensive$3 / $15
9GPT-5.6 Sol61.3OpenAIExtra high thinkingExtra high2.519721469—1482Expensive$4 / $20
10Gemini 3 Pro60.0Google (Gemini / DeepMind)Standard thinkingDefault——148414961481——
11Muse Spark 1.158.4Meta (Meta Superintelligence Labs, Muse)Standard thinkingDefault0.81927—1496—Mid-priced$1.25 / $4.25
12GPT-6 Astra57.7OpenAIMaximum thinkingMax3.52173144814891465Expensive$10 / $50
13GLM-5.356.6Z.ai (Zhipu)Maximum thinkingMax32075145514811469Mid-priced$1.40 / $4.40
14GPT-5.556.6OpenAIStandard thinkingDefault2.51844———Expensive$5 / $30
15Gemini 3.8 Flash56.2Google (Gemini / DeepMind)High thinkingHigh0.21748148415011486Mid-priced$0.75 / $3.75
16GPT-5.455.8OpenAIStandard thinkingDefault21840—1493—Mid-priced$2.50 / $15
17Gemini 3.7 Flash55.8Google (Gemini / DeepMind)High thinkingHigh-0.71723149614961498Mid-priced$0.75 / $3.75
18Muse Spark 1.255.5Meta (Meta Superintelligence Labs, Muse)Extra high thinkingExtra high-01840—1522—Mid-priced$1.25 / $4.25
19Muse Spark 1.354.6Meta (Meta Superintelligence Labs, Muse)Maximum thinkingMax0.61906145914911471Mid-priced$1.25 / $4.25
20Gemini 3.6 Flash53.4Google (Gemini / DeepMind)High thinkingHigh——1472—1475Mid-priced$0.75 / $3.75
21Claude Opus 4.852.6AnthropicHigh thinkingHigh0.81840146914981475Expensive$5 / $25
22Gemini 3.1 Pro (Preview)52.4Google (Gemini / DeepMind)Standard thinkingDefault-2.2—148014971480Mid-priced$2 / $12
23Claude Opus 4.550.3AnthropicHigh thinkingHigh——1469——Expensive$5 / $25
24Qwen3.8 Max (0803)47.3Qwen (Alibaba)Standard thinkingDefault——146814901472——
25Gemini 3.5 Flash47.1Google (Gemini / DeepMind)High thinkingHigh-2—1469—1473Mid-priced$1.50 / $9
26MiMo-V2.6-Pro45.7Xiaomi MiMoStandard thinkingDefault1.6—144514551457Cheap$0.43 / $0.87
27DeepSeek V4.1 Flash38.5DeepSeekMaximum thinkingMax-0.5—143814621451Cheap$0.30 / $1.20
28GPT-6 Sol36.5OpenAIMaximum thinkingMax—2125143814691440Mid-priced$2 / $10
29DeepSeek V4 Pro (0813)35.8DeepSeekHigh thinkingHigh0.5—144514721454Mid-priced$1.32 / $3.96
30Qwen3.8 27B34.3Qwen (Alibaba)Standard thinkingDefault-0.61671———Cheap$0.50 / $3
31Claude Sonnet 532.6AnthropicHigh thinkingHigh—1794143614741452Mid-priced$2 / $10
32Grok 4.729.2SpaceXAI (formerly xAI)Extra high thinkingExtra high0.42007143314251427Mid-priced$2 / $6
33GPT-6 Luna13.0OpenAIMaximum thinkingMax——140914551421Cheap$0.10 / $0.50

The score blends 9 tests for writing into one number out of 100: 100 is the best of all the models we track, 0 the weakest. “Best thinking level” is the setting where the model did best. The small dots show who ran each test: independent testers or the company that made the model. Price is cheap, mid-priced or expensive compared with the other models.

Composite = 0–100 weighted mean of 9 writing benchmarks, each min–max normalised across every model (or model × setting) that reports it. In the best-per-model view each benchmark uses the model's best setting (hover a value); “Mostly at” is the setting behind most of its score. Rows below the coverage threshold are dropped. Dots: independent, vendor-reported. Prices are list $/1M tokens.

Thinking levelReasoning effort

Does thinking longer help?Does more thinking help?Not always.

Many AI models let you choose how long they think before answering. Longer is slower and costs more, and past a point some models actually get worse. Pick a model and a test to see where it does best.Many models let you turn reasoning up. Past a point, some overthink: they spend more tokens and score lower. Pick a model and benchmark to see where it peaks.

Claude Opus 5.5 does best on this test at Medium. Letting it think longer, at Max, makes it slower and more expensive, and the result gets worse, not better.

Claude Opus 5.5 peaks at Medium on FrontierCode (54.6%). Turning it up to Max scores 54.4%, so more thinking costs tokens without buying accuracy.

Test by testBenchmark explorer

Who does bestWho leadson each test.on each benchmark.

Choose a test to see the fifteen models that did best on it. The colour shows who ran the test: independent testers, or the company that made the model.Top fifteen models per benchmark, each at its best setting. Colour shows who measured it: an independent leaderboard, or the vendor itself.

Measured by independent testersIndependent Reported by the makerVendor-reported

Short stories to constrained briefs, compared pairwise by a panel of LLM evaluators; Thurstone comparison score centered at 0.

Higher is better · each model at its best thinking level · 44 models tested · part of our writing & creativity score

Score · higher is better · best setting per model · 44 models scored · counts toward: Writing & creativity

Value for moneyPrice vs performance

What you getWhat a pointfor what you pay.of score costs.

Each dot is a model: higher means better results, further left means cheaper. The best deals sit up and to the left.Use-case composite against blended list price (3:1 input:output, log scale). Up and to the left is the most for the money.

Closed models (from a company)Proprietary Open models (can run on your own servers)Open weightsNamed: the best model you can get at each priceLabelled: Pareto front (nothing cheaper scores higher)

QuestionsFAQ

Things peopleCommonoften ask us.objections, answered.