Skip to content
Bencher

Models · OpenAI · Out sinceReleased 24 Feb 2026

GPT-5.3-Codex

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-5.3-Codex is made by OpenAI. We don't have enough test results yet to rank it. It's mid-priced to use.

Auto-created from OpenRouter catalog; verify details.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Mid-priced$1.75 / $14
Price per 1M (blended)Blended / 1M
$4.81
MemoryContext
400K
Longest answerMax output
—
Test resultsResults
8 (8 independent8 indep.)
Out sinceReleased
24 Feb 2026
Made byVendor
OpenAI
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.3-codex.

Route it as openai/gpt-5.3-codex at $1.75 in / $14 out per 1M tokens, 400K context. Listed since 24 Feb 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Epoch Capabilities Index156.8DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 153.5-160.8; best of listed model versions
IOI (Vals)53.8%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—±6.312 stderr; $2.681/test
LiveCodeBench87.3%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026—±0.968 stderr; $0.092/test
METR 50% time horizon5.8 hDefaultIndependent testIndependentMETR ↗8 May 2026METR react agent (Inspect)50% time horizon, METR-Horizon-v1.1; 95% CI 195-816 min; 80% horizon 54.7 min; release 2026-02-05; raw data https://metr.org/assets/benchmark_results_1_1.yaml
SWE Atlas - Codebase QnA32.6%Extra highGPT 5.3 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 7 (Scale rank accounts for CI); ±4.9; entry added 2026-02-25
SWE Atlas - Refactoring42.4%Extra highGPT 5.3 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-05-06
SWE Atlas - Test Writing39.0%Extra highGPT 5.3 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 2 (Scale rank accounts for CI); ±6.12; entry added 2026-03-26
SWE-bench Verified74.8%Highgpt-5.3-codex_highIndependent testIndependentEpoch AI ↗25 Feb 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 2.0pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)