Models · OpenAI · Out sinceReleased 5 Mar 2026
GPT-5.4#16 for writing.#16 for writing, best at its default setting.
GPT-5.4 is made by OpenAI. Among the models we track it ranks #16 for writing, #19 for coding. It's mid-priced to use.
Older reference point; prices/context from OpenRouter listing; scores only as listed in the GPT-5.5 launch post.
- 55.8 / 100 · #16
- —
- 51.1 / 100 · #19
- Mid-priced$2.50 / $15
- $5.63
- 1.1M
- 128K
- 53 (39 independent39 indep.)
- 5 Mar 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (its standard thinking level).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.4.
Route it as openai/gpt-5.4 at $2.50 in / $15 out per 1M tokens, 1.1M context. Listed since 5 Mar 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-5.4Where GPT-5.4think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
GSO
Best at Extra highBest at Extra high
LLM Creative Story-Writing Benchmark (Lech Mazur)
Best at Medium, worse abovePeaks at Medium
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| ARC-AGI-2 | 73.3% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | Verified; as listed in GPT-5.5 launch post |
| Artificial Analysis Intelligence Index | 39 | Extra highGPT-5.4 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-4; list price $2.5/15 per 1M in/out (AA marks this variant deprecated); estimated |
| Artificial Analysis output speed | 95 tok/s | Extra highGPT-5.4 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 188.6s; list price $2.5/15 per 1M in/out (AA marks this variant deprecated) |
| BrowseComp | 82.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | as listed in GPT-5.5 launch post |
| Epoch Capabilities Index | 156.9 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 154.9-159.1; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1840 | Defaultgpt-5.4 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 22; rubric score 16.89/20; slop 12.20; avg length 10488 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| EuroEval Swedish (generative) | 1.7 | No reasoningopenai/gpt-5.4-2026-03-05#none (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 6; ±0.09; lower is better; task scores (first metric): SweDN summarisation 35.65 ± 0.16, Skolprov 63.56 ± 2.78, Swedish facts 52.52 ± 2.85, ScaLA-sv 68.04 ± 2.02 |
| Expert-SWE (OpenAI internal) | 68.5% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | as listed in GPT-5.5 launch post |
| FrontierMath (Tiers 1-3) | 78.6% | Extra highgpt-5.4-2026-03-05_xhigh | Independent testIndependentEpoch AI ↗ | 11 Jun 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 27.1% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | as listed in GPT-5.5 launch post |
| GDPval | 83.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | wins or ties; as listed in GPT-5.5 launch post |
| GPQA Diamond | 93.3% | Extra highgpt-5.4-2026-03-05_xhigh | Independent testIndependentEpoch AI ↗ | 6 Mar 2026 | — | Epoch-run GPQA Diamond; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 92.0% | Extra highGPT-5.4 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-4) (AA marks this variant deprecated) |
| GPQA Diamond | 91.7% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.908 stderr; $0.209/test |
| GPQA Diamond | 92.8% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | as listed in GPT-5.5 launch post |
| GSO | 25.5% | Highreasoning_effort=high | Independent testIndependentGSO ↗ | 10 Mar 2026 | OpenHands | Opt@1; hack-controlled score 22.55; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable |
| GSO | 31.4% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | 10 Mar 2026 | OpenHands | Opt@1; hack-controlled score 30.39; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable |
| Humanity's Last Exam | 43.7% | Extra highGPT-5.4 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-4) (AA marks this variant deprecated) |
| Humanity's Last Exam | 36.2% | Extra highgpt-5.4-2026-03-05 (xhigh thinking) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 6 (Scale rank accounts for CI); ±1.88; entry added 2026-03-10 |
| Humanity's Last Exam | 39.8% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | no tools; as listed in GPT-5.5 launch post |
| Humanity's Last Exam (with tools) | 52.1% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | as listed in GPT-5.5 launch post |
| LegalBench (Vals) | 86.0% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-5.4-2026-03-05; rank 14/149; ±0.414 stderr; $0.023053/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 2 | MediumGPT-5.4 (medium) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 12/56; Thurstone comparison score (centered at 0); est. win chance 77%; 95% bootstrap 1.858 to 2.094 |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 2 | Extra highGPT-5.4 (xhigh) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 13/56; Thurstone comparison score (centered at 0); est. win chance 77%; 95% bootstrap 1.864 to 2.080 |
| LMArena Search Arena | 1197 | Defaultgpt-5.4-search | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 16 (rank range 10-18); 95% CI 1191.7-1202.6; 110116 votes |
| LMArena Text - Expert | 1515 | Highgpt-5.4-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 19 (rank range 8-50); 95% CI 1506.9-1523.6; 6334 votes; style-controlled |
| LMArena Text - Multi-Turn | 1493 | Highgpt-5.4-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 19 (rank range 7-44); 95% CI 1486.4-1499.7; 11980 votes; style-controlled |
| MCP Atlas | 70.6% | Extra highgpt-5.4 (xhigh) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 13 (Scale rank accounts for CI); ±2.8; entry added 2025-11-30 |
| MCP Atlas | 70.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | as listed in GPT-5.5 launch post |
| METR 50% time horizon | 5.7 h | Default | Independent testIndependentMETR ↗ | 8 May 2026 | METR react agent (Inspect) | 50% time horizon, METR-Horizon-v1.1; 95% CI 187-769 min; 80% horizon 53.9 min; release 2026-03-05; raw data https://metr.org/assets/benchmark_results_1_1.yaml |
| OSWorld-Verified | 75.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | as listed in GPT-5.5 launch post |
| PRBench Finance (Scale) | 45.6% | Highgpt-5.4 (High) | Independent testIndependentScale AI (SEAL) ↗ | 22 Apr 2026 | — | Scale rank 13; ±0.555 CI |
| PRBench Legal (Scale) | 44.4% | Highgpt-5.4 (High) | Independent testIndependentScale AI (SEAL) ↗ | 22 Apr 2026 | — | Scale rank 18; ±0.278 CI |
| ProgramBench (avg test pass rate) | 37.7% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 0.0%; avg cost $0.33/task |
| ProgramBench (fully resolved) | 0.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.5 stderr; $2.274/test; strict fully-resolved rate |
| ProgramBench (fully resolved) | 0.0% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 0.0%; avg cost $0.33/task |
| SimpleQA Verified | 45.1% | Extra highgpt-5.4-2026-03-05_xhigh | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.57 stderr |
| SkillsBench | 51.7% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.721 stderr; $2.044/test |
| SWE Atlas - Codebase QnA | 40.8% | Extra highGPT 5.4 (Codex) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 5 (Scale rank accounts for CI); ±5.1; entry added 2026-03-09 |
| SWE Atlas - Codebase QnA | 36.3% | Extra highGpt 5.4 xHigh (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 7 (Scale rank accounts for CI); ±4.9; entry added 2026-02-25 |
| SWE Atlas - Refactoring | 44.3% | Extra highGPT 5.4 (Codex) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-05-06 |
| SWE Atlas - Test Writing | 44.4% | Extra highGPT-5.4 (Codex) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 1 (Scale rank accounts for CI); ±6.04; entry added 2026-03-26 |
| SWE Atlas - Test Writing | 40.0% | Extra highGPT-5.4 (Mini-SWE) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 2 (Scale rank accounts for CI); ±6; entry added 2026-03-26 |
| SWE-Bench Pro (private/commercial set) | 43.4% | Extra highgpt-5.4(xHigh)* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±6.03; entry added 2026-04-08; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 59.1% | Extra highgpt-5.4 (xHigh)* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 1 (Scale rank accounts for CI); ±3.56; entry added 2026-04-08; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 57.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | public set; as listed in GPT-5.5 launch post |
| SWE-bench Verified | 76.9% | Highgpt-5.4-2026-03-05_high | Independent testIndependentEpoch AI ↗ | 6 Mar 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| tau2-bench | 87.1% | Extra highGPT-5.4 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-4) (AA marks this variant deprecated) |
| Terminal-Bench 2.0 | 75.1% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | Terminal-Bench 2.0; as listed in GPT-5.5 launch post |
| Terminal-Bench 2.1 | 78.3% | Extra highGPT-5.4 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-4) (AA marks this variant deprecated) |
| Toolathlon | 54.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | as listed in GPT-5.5 launch post |
| Vectara Hallucination Leaderboard (HHEM) | 7.0% | Defaultopenai/gpt-5.4-2026-03-05 | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 93.0 %; answer rate 99.9 %; avg summary 81.7 words; HHEM-2.3 judge; effort not stated (API default) |
| τ²-bench (Telecom) | 92.8% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | original prompts; as listed in GPT-5.5 launch post |