Skip to content
Bencher

Models · OpenAI · Out sinceReleased 5 Mar 2026

GPT-5.4 Pro

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-5.4 Pro is made by OpenAI. We don't have enough test results yet to rank it. It's expensive to use.

Auto-created from OpenRouter catalog; verify details.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Expensive$30 / $180
Price per 1M (blended)Blended / 1M
$67.50
MemoryContext
1.1M
Longest answerMax output
—
Test resultsResults
7 (7 independent7 indep.)
Out sinceReleased
5 Mar 2026
Made byVendor
OpenAI
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.4-pro.

Route it as openai/gpt-5.4-pro at $30 in / $180 out per 1M tokens, 1.1M context. Listed since 5 Mar 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Epoch Capabilities Index159.1DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 156.5-161.9; best of listed model versions
FrontierMath (Tiers 1-3)82.5%Extra highgpt-5.4-pro-2026-03-05_xhighIndependent testIndependentEpoch AI ↗13 Jun 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.3pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 458.5%Extra highgpt-5.4-pro-2026-03-05_xhighIndependent testIndependentEpoch AI ↗13 Jun 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 7.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
GPQA Diamond94.6%Extra highgpt-5.4-pro-2026-03-05_xhighIndependent testIndependentEpoch AI ↗20 Mar 2026—Epoch-run GPQA Diamond; stderr 1.6pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
Humanity's Last Exam44.3%Defaultgpt-5.4-pro-2026-03-05Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±1.95; entry added 2026-03-23
SimpleQA Verified46.3%Extra highgpt-5.4-pro-2026-03-05_xhighIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
Vectara Hallucination Leaderboard (HHEM)8.3%Defaultopenai/gpt-5.4-pro-2026-03-05Independent testIndependentVectara ↗22 Sep 2026—factual consistency 91.7 %; answer rate 100.0 %; avg summary 148.5 words; HHEM-2.3 judge; effort not stated (API default)