Skip to content
Bencher

Models · OpenAI · Out sinceReleased 5 Mar 2026

GPT-5.4#16 for writing.#16 for writing, best at its default setting.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-5.4 is made by OpenAI. Among the models we track it ranks #16 for writing, #19 for coding. It's mid-priced to use.

Older reference point; prices/context from OpenRouter listing; scores only as listed in the GPT-5.5 launch post.

Writing & creativity
55.8 / 100 · #16
Research & analysis
—
Coding
51.1 / 100 · #19
Price
Mid-priced$2.50 / $15
Price per 1M (blended)Blended / 1M
$5.63
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
53 (39 independent39 indep.)
Out sinceReleased
5 Mar 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowMediumHighExtra high

Highlighted: where it did best for writing (its standard thinking level).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.4.

Route it as openai/gpt-5.4 at $2.50 in / $15 out per 1M tokens, 1.1M context. Listed since 5 Mar 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-5.4Where GPT-5.4think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

GSO

Best at Extra highBest at Extra high

LLM Creative Story-Writing Benchmark (Lech Mazur)

Best at Medium, worse abovePeaks at Medium

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
ARC-AGI-273.3%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—Verified; as listed in GPT-5.5 launch post
Artificial Analysis Intelligence Index39Extra highGPT-5.4 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-4; list price $2.5/15 per 1M in/out (AA marks this variant deprecated); estimated
Artificial Analysis output speed95 tok/sExtra highGPT-5.4 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 188.6s; list price $2.5/15 per 1M in/out (AA marks this variant deprecated)
BrowseComp82.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—as listed in GPT-5.5 launch post
Epoch Capabilities Index156.9DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 154.9-159.1; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1840Defaultgpt-5.4Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 22; rubric score 16.89/20; slop 12.20; avg length 10488 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
EuroEval Swedish (generative)1.7No reasoningopenai/gpt-5.4-2026-03-05#none (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 6; ±0.09; lower is better; task scores (first metric): SweDN summarisation 35.65 ± 0.16, Skolprov 63.56 ± 2.78, Swedish facts 52.52 ± 2.85, ScaLA-sv 68.04 ± 2.02
Expert-SWE (OpenAI internal)68.5%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—as listed in GPT-5.5 launch post
FrontierMath (Tiers 1-3)78.6%Extra highgpt-5.4-2026-03-05_xhighIndependent testIndependentEpoch AI ↗11 Jun 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 427.1%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—as listed in GPT-5.5 launch post
GDPval83.0%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—wins or ties; as listed in GPT-5.5 launch post
GPQA Diamond93.3%Extra highgpt-5.4-2026-03-05_xhighIndependent testIndependentEpoch AI ↗6 Mar 2026—Epoch-run GPQA Diamond; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond92.0%Extra highGPT-5.4 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-4) (AA marks this variant deprecated)
GPQA Diamond91.7%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026—±1.908 stderr; $0.209/test
GPQA Diamond92.8%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—as listed in GPT-5.5 launch post
GSO25.5%Highreasoning_effort=highIndependent testIndependentGSO ↗10 Mar 2026OpenHandsOpt@1; hack-controlled score 22.55; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
GSO31.4%Extra highreasoning_effort=xhighIndependent testIndependentGSO ↗10 Mar 2026OpenHandsOpt@1; hack-controlled score 30.39; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
Humanity's Last Exam43.7%Extra highGPT-5.4 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-4) (AA marks this variant deprecated)
Humanity's Last Exam36.2%Extra highgpt-5.4-2026-03-05 (xhigh thinking)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 6 (Scale rank accounts for CI); ±1.88; entry added 2026-03-10
Humanity's Last Exam39.8%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—no tools; as listed in GPT-5.5 launch post
Humanity's Last Exam (with tools)52.1%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—as listed in GPT-5.5 launch post
LegalBench (Vals)86.0%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-5.4-2026-03-05; rank 14/149; ±0.414 stderr; $0.023053/test
LLM Creative Story-Writing Benchmark (Lech Mazur)2MediumGPT-5.4 (medium)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 12/56; Thurstone comparison score (centered at 0); est. win chance 77%; 95% bootstrap 1.858 to 2.094
LLM Creative Story-Writing Benchmark (Lech Mazur)2Extra highGPT-5.4 (xhigh)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 13/56; Thurstone comparison score (centered at 0); est. win chance 77%; 95% bootstrap 1.864 to 2.080
LMArena Search Arena1197Defaultgpt-5.4-searchIndependent testIndependentLMArena ↗24 Aug 2026—rank 16 (rank range 10-18); 95% CI 1191.7-1202.6; 110116 votes
LMArena Text - Expert1515Highgpt-5.4-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 8-50); 95% CI 1506.9-1523.6; 6334 votes; style-controlled
LMArena Text - Multi-Turn1493Highgpt-5.4-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 7-44); 95% CI 1486.4-1499.7; 11980 votes; style-controlled
MCP Atlas70.6%Extra highgpt-5.4 (xhigh)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 13 (Scale rank accounts for CI); ±2.8; entry added 2025-11-30
MCP Atlas70.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—as listed in GPT-5.5 launch post
METR 50% time horizon5.7 hDefaultIndependent testIndependentMETR ↗8 May 2026METR react agent (Inspect)50% time horizon, METR-Horizon-v1.1; 95% CI 187-769 min; 80% horizon 53.9 min; release 2026-03-05; raw data https://metr.org/assets/benchmark_results_1_1.yaml
OSWorld-Verified75.0%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—as listed in GPT-5.5 launch post
PRBench Finance (Scale)45.6%Highgpt-5.4 (High)Independent testIndependentScale AI (SEAL) ↗22 Apr 2026—Scale rank 13; ±0.555 CI
PRBench Legal (Scale)44.4%Highgpt-5.4 (High)Independent testIndependentScale AI (SEAL) ↗22 Apr 2026—Scale rank 18; ±0.278 CI
ProgramBench (avg test pass rate)37.7%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 0.0%; avg cost $0.33/task
ProgramBench (fully resolved)0.5%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.5 stderr; $2.274/test; strict fully-resolved rate
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 0.0%; avg cost $0.33/task
SimpleQA Verified45.1%Extra highgpt-5.4-2026-03-05_xhighIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.57 stderr
SkillsBench51.7%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.721 stderr; $2.044/test
SWE Atlas - Codebase QnA40.8%Extra highGPT 5.4 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 5 (Scale rank accounts for CI); ±5.1; entry added 2026-03-09
SWE Atlas - Codebase QnA36.3%Extra highGpt 5.4 xHigh (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 7 (Scale rank accounts for CI); ±4.9; entry added 2026-02-25
SWE Atlas - Refactoring44.3%Extra highGPT 5.4 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-05-06
SWE Atlas - Test Writing44.4%Extra highGPT-5.4 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 1 (Scale rank accounts for CI); ±6.04; entry added 2026-03-26
SWE Atlas - Test Writing40.0%Extra highGPT-5.4 (Mini-SWE) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 2 (Scale rank accounts for CI); ±6; entry added 2026-03-26
SWE-Bench Pro (private/commercial set)43.4%Extra highgpt-5.4(xHigh)*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±6.03; entry added 2026-04-08; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)59.1%Extra highgpt-5.4 (xHigh)*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 1 (Scale rank accounts for CI); ±3.56; entry added 2026-04-08; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)57.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—public set; as listed in GPT-5.5 launch post
SWE-bench Verified76.9%Highgpt-5.4-2026-03-05_highIndependent testIndependentEpoch AI ↗6 Mar 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)
tau2-bench87.1%Extra highGPT-5.4 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-4) (AA marks this variant deprecated)
Terminal-Bench 2.075.1%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—Terminal-Bench 2.0; as listed in GPT-5.5 launch post
Terminal-Bench 2.178.3%Extra highGPT-5.4 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-4) (AA marks this variant deprecated)
Toolathlon54.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—as listed in GPT-5.5 launch post
Vectara Hallucination Leaderboard (HHEM)7.0%Defaultopenai/gpt-5.4-2026-03-05Independent testIndependentVectara ↗22 Sep 2026—factual consistency 93.0 %; answer rate 99.9 %; avg summary 81.7 words; HHEM-2.3 judge; effort not stated (API default)
τ²-bench (Telecom)92.8%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—original prompts; as listed in GPT-5.5 launch post