Skip to content
Bencher

Models · OpenAI · Out sinceReleased 17 Mar 2026

GPT-5.4 Mini

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-5.4 Mini is made by OpenAI. We don't have enough test results yet to rank it. It's mid-priced to use.

Auto-created from OpenRouter catalog; verify details.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Mid-priced$0.75 / $4.50
Price per 1M (blended)Blended / 1M
$1.69
MemoryContext
400K
Longest answerMax output
—
Test resultsResults
6 (6 independent6 indep.)
Out sinceReleased
17 Mar 2026
Made byVendor
OpenAI
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.4-mini.

Route it as openai/gpt-5.4-mini at $0.75 in / $4.50 out per 1M tokens, 400K context. Listed since 17 Mar 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-5.4 MiniWhere GPT-5.4 Minithink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

EuroEval Swedish (generative)

Best at HighBest at High

Lower is better on this test

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
EuroEval Swedish (generative)1.7Lowopenai/gpt-5.4-mini-2026-03-17#low (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 5; ±0.07; lower is better; task scores (first metric): SweDN summarisation 36.78 ± 0.17, Skolprov 75.69 ± 3.17, Swedish facts 68.19 ± 2.71, ScaLA-sv 67.25 ± 1.40
EuroEval Swedish (generative)1.6Mediumopenai/gpt-5.4-mini-2026-03-17#medium (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 4; ±0.08; lower is better; task scores (first metric): SweDN summarisation 36.17 ± 0.17, Skolprov 76.36 ± 2.15, Swedish facts 65.91 ± 2.71, ScaLA-sv 70.19 ± 1.35
EuroEval Swedish (generative)1.5Highopenai/gpt-5.4-mini-2026-03-17#high (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 3; ±0.09; lower is better; task scores (first metric): SweDN summarisation 37.23 ± 0.19, Skolprov 86.03 ± 1.89, Swedish facts 68.70 ± 2.30, ScaLA-sv 67.69 ± 1.31
ProgramBench (avg test pass rate)16.4%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 0.0%; avg cost $0.04/task
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 0.0%; avg cost $0.04/task
Vectara Hallucination Leaderboard (HHEM)5.5%Defaultopenai/gpt-5.4-mini-2026-03-17Independent testIndependentVectara ↗22 Sep 2026—factual consistency 94.5 %; answer rate 100.0 %; avg summary 54.7 words; HHEM-2.3 judge; effort not stated (API default)