Models · OpenAI · Out sinceReleased 23 Apr 2026
GPT-5.5#14 for writing.#14 for writing, best at its default setting.
GPT-5.5 is made by OpenAI. Among the models we track it ranks #14 for writing, #16 for coding, #24 for research and analysis. It's expensive to use.
Snapshot gpt-5.5-2026-04-23 (API ~Apr 24). Default effort medium. Retiring from ChatGPT/Codex Oct 14, 2026 (API unaffected). Launch evals run at xhigh.
- 56.6 / 100 · #14
- 61.5 / 100 · #24
- 52.8 / 100 · #16
- Expensive$5 / $30
- $11.25
- 1.1M
- 128K
- 113 (70 independent70 indep.)
- 23 Apr 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (its standard thinking level).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.5.
Route it as openai/gpt-5.5 at $5 in / $30 out per 1M tokens, 1.1M context. Listed since 24 Apr 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-5.5Where GPT-5.5think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at Extra highBest at Extra high
DeepSWE
Best at Extra highBest at Extra high
ProgramBench (avg test pass rate)
Best at Extra highBest at Extra high
ProgramBench (fully resolved)
Best at HighBest at High
Terminal-Bench 2.1
Best at Extra highBest at Extra high
SciCode
Best at High, worse abovePeaks at High
Artificial Analysis Intelligence Index
Best at Extra highBest at Extra high
tau2-bench
Best at Extra highBest at Extra high
ARC-AGI-2
Best at Extra highBest at Extra high
GPQA Diamond
Best at Extra highBest at Extra high
Humanity's Last Exam
Best at Extra highBest at Extra high
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 50.0% | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| Agents' Last Exam | 46.9% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| ARC-AGI-1 | 95.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | Verified |
| ARC-AGI-2 | 83.3% | HighGPT-5.5 (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.45/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 85.0% | Extra highGPT-5.5 (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.87/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 85.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | Verified |
| ARC-AGI-3 | 0.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| Artificial Analysis Coding Agent Index v1.1 | 76.4 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | Artificial Analysis Coding Agent Index v1.1 (index score) |
| Artificial Analysis Intelligence Index | 33.8 | MediumGPT-5.5 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-5-medium; list price $5/30 per 1M in/out; cost to run AA Intelligence Index $0.90/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 37 | HighGPT-5.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-5-high; list price $5/30 per 1M in/out; cost to run AA Intelligence Index $1.54/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 38.4 | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-5; list price $5/30 per 1M in/out; cost to run AA Intelligence Index $2.63/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 54.8 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | Artificial Analysis Intelligence Index v4.1 |
| Artificial Analysis output speed | 82 tok/s | MediumGPT-5.5 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 7.5s; list price $5/30 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 84 tok/s | HighGPT-5.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 26.2s; list price $5/30 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 86 tok/s | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 48.1s; list price $5/30 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 12.9% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| BrowseComp | 84.4% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| BrowseComp | 84.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| DeepSWE | 54.0% | Mediumgpt-5.5_medium | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 77.9%; ±2.6; 4 runs; $2.75/task |
| DeepSWE | 64.4% | Highgpt-5.5_high | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 90.3%; ±3.1; 4 runs; $5.10/task |
| DeepSWE | 67.0% | Extra highgpt-5.5_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 88.5%; ±6.5; 4 runs; $7.23/task |
| DeepSWE | 67.0% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | |
| Epoch Capabilities Index | 159.2 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 156.9-162.4; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1844 | Defaultgpt-5.5 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 18; rubric score 17.01/20; slop 13.10; avg length 12945 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| Expert-SWE (OpenAI internal) | 73.1% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | OpenAI internal long-horizon coding eval |
| Finance Agent | 60.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | FinanceAgent v1.1 |
| FrontierCode | 43.0% | Defaultgpt-5.5_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | codex | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5 |
| FrontierMath (Tiers 1-3) | 85.3% | Extra highgpt-5.5_xhigh | Independent testIndependentEpoch AI ↗ | 11 Jun 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 1-3 | 51.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| FrontierMath Tier 1-3 | 85.3% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| FrontierMath Tier 4 | 72.5% | Extra highgpt-5.5_xhigh | Independent testIndependentEpoch AI ↗ | 11 Jun 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 7.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 35.4% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| FrontierMath Tier 4 | 72.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GDP.pdf | 26.0% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GDPval | 84.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | wins or ties |
| GDPval-AA v2 | 1494 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GPQA Diamond | 92.6% | MediumGPT-5.5 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated) |
| GPQA Diamond | 93.2% | HighGPT-5.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated) |
| GPQA Diamond | 94.0% | Extra highgpt-5.5-pre-release_xhigh | Independent testIndependentEpoch AI ↗ | 24 Apr 2026 | — | Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv); pre-release checkpoint |
| GPQA Diamond | 93.5% | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated) |
| GPQA Diamond | 93.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.288 stderr; $0.083/test |
| GPQA Diamond | 93.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| GPQA Diamond | 93.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GSO | 40.2% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | 27 Apr 2026 | OpenHands | Opt@1; hack-controlled score 37.25; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable |
| HealthBench Professional | 49.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | official HealthBench Professional scoring |
| Humanity's Last Exam | 42.4% | MediumGPT-5.5 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated) |
| Humanity's Last Exam | 45.0% | HighGPT-5.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated) |
| Humanity's Last Exam | 45.8% | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated) |
| Humanity's Last Exam | 41.4% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | no tools |
| Humanity's Last Exam (with tools) | 52.2% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| LegalBench (Vals) | 86.5% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-5.5; rank 12/149; ±0.41 stderr; $0.015403/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 2.5 | Extra highGPT-5.5 (xhigh) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 10/56; Thurstone comparison score (centered at 0); est. win chance 82%; 95% bootstrap 2.386 to 2.522 |
| LMArena Search Arena | 1242 | Defaultgpt-5.5-search | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 3 (rank range 3-5); 95% CI 1236.7-1247.4; 89873 votes |
| LMArena Text - Occupational: Business, Management & Finance | 1490 | Highgpt-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 11 (rank range 3-31); 95% CI 1483.2-1495.9; 13729 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1486 | Defaultgpt-5.5 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 7-37); 95% CI 1479.2-1491.8; 14079 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1495 | Highgpt-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 3-50); 95% CI 1486.0-1503.6; 5614 votes; style-controlled |
| LMArena Text (overall) | 1481 | Highgpt-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 22 (CI rank 13-39); 95% CI 1478-1485; 69312 votes |
| LMArena Text (overall) | 1477 | Defaultgpt-5.5 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 28 (CI rank 20-44); 95% CI 1473-1480; 70894 votes |
| MCP Atlas | 75.3% | Extra highgpt-5.5 (xhigh) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±2.7; entry added 2026-04-28 |
| MCP Atlas | 75.3% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| MMMU-Pro (no tools) | 81.2% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | no tools |
| MMMU-Pro (no tools) | 81.2% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| MMMU-Pro (with tools) | 83.2% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| MMMU-Pro (with tools) | 83.2% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OfficeQA Pro | 54.1% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 81.5% | Extra highreasoning effort=xhigh; 8-needle 256K-512K | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 74.0% | Extra highreasoning effort=xhigh; 8-needle 512K-1M | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 81.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 256K-512K | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 74.0% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 512K-1M | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OSWorld 2.0 | 13.0% | Extra highgpt-5.5_xhigh | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.495; tool setting batch tool; step budget 500 |
| OSWorld 2.0 | 47.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | OSWorld 2.0 (OpenAI-run) |
| OSWorld-Verified | 78.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| ProgramBench (avg test pass rate) | 66.5% | High | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.5%, almost (>=95% tests) 5.0%; avg cost $3.65/task |
| ProgramBench (avg test pass rate) | 69.5% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.5%, almost (>=95% tests) 13.5%; avg cost $8.85/task |
| ProgramBench (avg test pass rate) | 56.6% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 1.5%; avg cost $1.21/task |
| ProgramBench (fully resolved) | 0.5% | High | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 5.0%; avg cost $3.65/task |
| ProgramBench (fully resolved) | 0.5% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.5 stderr; $6.955/test; strict fully-resolved rate |
| ProgramBench (fully resolved) | 0.5% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 13.5%; avg cost $8.85/task |
| ProgramBench (fully resolved) | 0.0% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 1.5%; avg cost $1.21/task |
| SciCode | 54.5% | MediumGPT-5.5 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated) |
| SciCode | 56.1% | HighGPT-5.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated) |
| SciCode | 55.8% | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated) |
| SimpleBench | 69.0% | DefaultGPT-5.5 | Independent testIndependentSimpleBench ↗ | 24 Apr 2026 | — | AVG@5, temp 0.7; rank 19th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 63.0% | Extra highgpt-5.5_xhigh | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.53 stderr |
| SkillsBench | 62.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.523 stderr; $2.543/test |
| SWE Atlas - Codebase QnA | 45.4% | Extra highGPT 5.5 (Codex) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 2 (Scale rank accounts for CI); ±5.08; entry added 2026-05-07 |
| SWE Atlas - Refactoring | 44.8% | Extra highGPT-5.5 (Codex) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-05-06 |
| SWE Atlas - Test Writing | 42.6% | Extra highGPT-5.5 (Codex) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 2 (Scale rank accounts for CI); ±5.96; entry added 2026-05-07 |
| SWE-Bench Pro (public, v1) | 58.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | public set |
| SWE-Bench Pro (public, v1) | 59.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| SWE-bench Verified | 82.6% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.695 stderr; $1.362/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| SWE-bench Verified | 80.6% | Extra highgpt-5.5-pre-release_xhigh | Independent testIndependentEpoch AI ↗ | 24 Apr 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv); pre-release checkpoint |
| tau2-bench | 91.8% | MediumGPT-5.5 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated) |
| tau2-bench | 93.0% | HighGPT-5.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated) |
| tau2-bench | 93.9% | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated) |
| Terminal-Bench 2.0 | 82.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | Codex | Terminal-Bench 2.0 |
| Terminal-Bench 2.1 | 80.5% | MediumGPT-5.5 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 79.4% | HighGPT-5.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 76.4% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.649 stderr; $0.742/test |
| Terminal-Bench 2.1 | 84.3% | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 83.1% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Codex CLI | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 78.0% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Terminus 2 | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 85.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | Terminal-Bench 2.1 |
| Terminal-Bench 4.0 | 5.1% | MediumGPT-5.5 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 9.1% | HighGPT-5.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 14.6% | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated) |
| Toolathlon | 55.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | |
| Toolathlon | 55.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| Vals Code Migration | 45.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.164 stderr; $6.442/test |
| Vals CorpFin v2 | 68.4% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id openai/gpt-5.5; rank 9/134; ±0.916 stderr; $0.43166/test |
| Vals SRE Bench | 3.8% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±1.186 stderr; $17.774/test |
| Vectara Hallucination Leaderboard (HHEM) | 9.3% | Defaultopenai/gpt-5.5 | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 90.7 %; answer rate 100.0 %; avg summary 129.6 words; HHEM-2.3 judge; effort not stated (API default) |
| τ²-bench (Telecom) | 98.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 23 Apr 2026 | — | original prompts, no prompt tuning, GPT-4.1 user model |