Models · OpenAI · Out sinceReleased 10 Dec 2025
GPT-5.2
GPT-5.2 is made by OpenAI. We don't have enough test results yet to rank it. It's mid-priced to use.
Auto-created from OpenRouter catalog; verify details.
- —
- —
- —
- Mid-priced$1.75 / $14
- $4.81
- 400K
- —
- 9 (9 independent9 indep.)
- 10 Dec 2025
- OpenAI
- —
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.2.
Route it as openai/gpt-5.2 at $1.75 in / $14 out per 1M tokens, 400K context. Listed since 10 Dec 2025.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| GPQA Diamond | 91.7% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.837 stderr; $0.149/test |
| LMArena Search Arena | 1207 | Defaultgpt-5.2-search | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 11 (rank range 8-16); 95% CI 1201.3-1212.5; 52712 votes |
| MCP Atlas | 67.6% | Extra highgpt-5.2 (xhigh) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 17 (Scale rank accounts for CI); ±2.9; entry added 2025-09-10 |
| METR 50% time horizon | 5.9 h | Default | Independent testIndependentMETR ↗ | 8 May 2026 | METR react agent (Inspect) | 50% time horizon, METR-Horizon-v1.1; 95% CI 198-815 min; 80% horizon 66.0 min; release 2025-12-11; raw data https://metr.org/assets/benchmark_results_1_1.yaml |
| SWE-bench Multilingual | 66.7% | High | Independent testIndependentSWE-bench ↗ | 13 Feb 2026 | mini-SWE-agent 2.0.0a0 | Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26 |
| SWE-Bench Pro (private/commercial set) | 23.8% | DefaultGPT 5.2 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 7 (Scale rank accounts for CI); ±5.09; entry added 2026-01-12 |
| SWE-Bench Pro (public, v1) | 29.9% | Defaultgpt-5.2 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 16 (Scale rank accounts for CI); ±2.15; entry added 2026-01-27 |
| SWE-bench Verified (bash-only, mini-SWE-agent) | 72.8% | High | Independent testIndependentSWE-bench ↗ | 17 Feb 2026 | mini-SWE-agent 2.0.0 | Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26 |
| Vals TaxEval v2 | 75.8% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id openai/gpt-5.2-2025-12-11; rank 11/145; ±0.845 stderr; $0.066435/test |