Skip to content
Bencher

Models · OpenAI · Out sinceReleased 23 Apr 2026

GPT-5.5#14 for writing.#14 for writing, best at its default setting.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary
Where you can use it:In apps:ChatGPT (Business plan) · GPT-5.5

GPT-5.5 is made by OpenAI. Among the models we track it ranks #14 for writing, #16 for coding, #24 for research and analysis. It's expensive to use.

Snapshot gpt-5.5-2026-04-23 (API ~Apr 24). Default effort medium. Retiring from ChatGPT/Codex Oct 14, 2026 (API unaffected). Launch evals run at xhigh.

Writing & creativity
56.6 / 100 · #14
Research & analysis
61.5 / 100 · #24
Coding
52.8 / 100 · #16
Price
Expensive$5 / $30
Price per 1M (blended)Blended / 1M
$11.25
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
113 (70 independent70 indep.)
Out sinceReleased
23 Apr 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowMediumHighExtra high

Highlighted: where it did best for writing (its standard thinking level).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.5.

Route it as openai/gpt-5.5 at $5 in / $30 out per 1M tokens, 1.1M context. Listed since 24 Apr 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-5.5Where GPT-5.5think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at Extra highBest at Extra high

DeepSWE

Best at Extra highBest at Extra high

ProgramBench (avg test pass rate)

Best at Extra highBest at Extra high

ProgramBench (fully resolved)

Best at HighBest at High

Terminal-Bench 2.1

Best at Extra highBest at Extra high

SciCode

Best at High, worse abovePeaks at High

Artificial Analysis Intelligence Index

Best at Extra highBest at Extra high

tau2-bench

Best at Extra highBest at Extra high

ARC-AGI-2

Best at Extra highBest at Extra high

GPQA Diamond

Best at Extra highBest at Extra high

Humanity's Last Exam

Best at Extra highBest at Extra high

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent50.0%Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
Agents' Last Exam46.9%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
ARC-AGI-195.0%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—Verified
ARC-AGI-283.3%HighGPT-5.5 (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.45/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-285.0%Extra highGPT-5.5 (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.87/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-285.0%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—Verified
ARC-AGI-30.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
Artificial Analysis Coding Agent Index v1.176.4DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—Artificial Analysis Coding Agent Index v1.1 (index score)
Artificial Analysis Intelligence Index33.8MediumGPT-5.5 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-5-medium; list price $5/30 per 1M in/out; cost to run AA Intelligence Index $0.90/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index37HighGPT-5.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-5-high; list price $5/30 per 1M in/out; cost to run AA Intelligence Index $1.54/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index38.4Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-5; list price $5/30 per 1M in/out; cost to run AA Intelligence Index $2.63/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index54.8DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—Artificial Analysis Intelligence Index v4.1
Artificial Analysis output speed82 tok/sMediumGPT-5.5 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 7.5s; list price $5/30 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed84 tok/sHighGPT-5.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 26.2s; list price $5/30 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed86 tok/sExtra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 48.1s; list price $5/30 per 1M in/out (AA marks this variant deprecated)
AutomationBench12.9%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
BrowseComp84.4%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
BrowseComp84.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
DeepSWE54.0%Mediumgpt-5.5_mediumIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 77.9%; ±2.6; 4 runs; $2.75/task
DeepSWE64.4%Highgpt-5.5_highIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 90.3%; ±3.1; 4 runs; $5.10/task
DeepSWE67.0%Extra highgpt-5.5_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 88.5%; ±6.5; 4 runs; $7.23/task
DeepSWE67.0%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026Codex
Epoch Capabilities Index159.2DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 156.9-162.4; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1844Defaultgpt-5.5Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 18; rubric score 17.01/20; slop 13.10; avg length 12945 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
Expert-SWE (OpenAI internal)73.1%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—OpenAI internal long-horizon coding eval
Finance Agent60.0%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—FinanceAgent v1.1
FrontierCode43.0%Defaultgpt-5.5_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026codexread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5
FrontierMath (Tiers 1-3)85.3%Extra highgpt-5.5_xhighIndependent testIndependentEpoch AI ↗11 Jun 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 1-351.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
FrontierMath Tier 1-385.3%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
FrontierMath Tier 472.5%Extra highgpt-5.5_xhighIndependent testIndependentEpoch AI ↗11 Jun 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 7.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 435.4%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
FrontierMath Tier 472.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GDP.pdf26.0%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GDPval84.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—wins or ties
GDPval-AA v21494DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GPQA Diamond92.6%MediumGPT-5.5 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated)
GPQA Diamond93.2%HighGPT-5.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated)
GPQA Diamond94.0%Extra highgpt-5.5-pre-release_xhighIndependent testIndependentEpoch AI ↗24 Apr 2026—Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv); pre-release checkpoint
GPQA Diamond93.5%Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated)
GPQA Diamond93.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026—±1.288 stderr; $0.083/test
GPQA Diamond93.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
GPQA Diamond93.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GSO40.2%Extra highreasoning_effort=xhighIndependent testIndependentGSO ↗27 Apr 2026OpenHandsOpt@1; hack-controlled score 37.25; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
HealthBench Professional49.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—official HealthBench Professional scoring
Humanity's Last Exam42.4%MediumGPT-5.5 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated)
Humanity's Last Exam45.0%HighGPT-5.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated)
Humanity's Last Exam45.8%Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated)
Humanity's Last Exam41.4%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—no tools
Humanity's Last Exam (with tools)52.2%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
LegalBench (Vals)86.5%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-5.5; rank 12/149; ±0.41 stderr; $0.015403/test
LLM Creative Story-Writing Benchmark (Lech Mazur)2.5Extra highGPT-5.5 (xhigh)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 10/56; Thurstone comparison score (centered at 0); est. win chance 82%; 95% bootstrap 2.386 to 2.522
LMArena Search Arena1242Defaultgpt-5.5-searchIndependent testIndependentLMArena ↗24 Aug 2026—rank 3 (rank range 3-5); 95% CI 1236.7-1247.4; 89873 votes
LMArena Text - Occupational: Business, Management & Finance1490Highgpt-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 11 (rank range 3-31); 95% CI 1483.2-1495.9; 13729 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1486Defaultgpt-5.5Independent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 7-37); 95% CI 1479.2-1491.8; 14079 votes; style-controlled
LMArena Text - Occupational: Legal & Government1495Highgpt-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 3-50); 95% CI 1486.0-1503.6; 5614 votes; style-controlled
LMArena Text (overall)1481Highgpt-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 22 (CI rank 13-39); 95% CI 1478-1485; 69312 votes
LMArena Text (overall)1477Defaultgpt-5.5Independent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 28 (CI rank 20-44); 95% CI 1473-1480; 70894 votes
MCP Atlas75.3%Extra highgpt-5.5 (xhigh)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±2.7; entry added 2026-04-28
MCP Atlas75.3%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
MMMU-Pro (no tools)81.2%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—no tools
MMMU-Pro (no tools)81.2%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
MMMU-Pro (with tools)83.2%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
MMMU-Pro (with tools)83.2%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OfficeQA Pro54.1%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
OpenAI MRCR v2 (8-needle)81.5%Extra highreasoning effort=xhigh; 8-needle 256K-512KMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
OpenAI MRCR v2 (8-needle)74.0%Extra highreasoning effort=xhigh; 8-needle 512K-1MMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
OpenAI MRCR v2 (8-needle)81.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 256K-512KMaker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OpenAI MRCR v2 (8-needle)74.0%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 512K-1MMaker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OSWorld 2.013.0%Extra highgpt-5.5_xhighIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.495; tool setting batch tool; step budget 500
OSWorld 2.047.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—OSWorld 2.0 (OpenAI-run)
OSWorld-Verified78.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
ProgramBench (avg test pass rate)66.5%HighIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.5%, almost (>=95% tests) 5.0%; avg cost $3.65/task
ProgramBench (avg test pass rate)69.5%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.5%, almost (>=95% tests) 13.5%; avg cost $8.85/task
ProgramBench (avg test pass rate)56.6%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 1.5%; avg cost $1.21/task
ProgramBench (fully resolved)0.5%HighIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 5.0%; avg cost $3.65/task
ProgramBench (fully resolved)0.5%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.5 stderr; $6.955/test; strict fully-resolved rate
ProgramBench (fully resolved)0.5%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 13.5%; avg cost $8.85/task
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 1.5%; avg cost $1.21/task
SciCode54.5%MediumGPT-5.5 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated)
SciCode56.1%HighGPT-5.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated)
SciCode55.8%Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated)
SimpleBench69.0%DefaultGPT-5.5Independent testIndependentSimpleBench ↗24 Apr 2026—AVG@5, temp 0.7; rank 19th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified63.0%Extra highgpt-5.5_xhighIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.53 stderr
SkillsBench62.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.523 stderr; $2.543/test
SWE Atlas - Codebase QnA45.4%Extra highGPT 5.5 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 2 (Scale rank accounts for CI); ±5.08; entry added 2026-05-07
SWE Atlas - Refactoring44.8%Extra highGPT-5.5 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-05-06
SWE Atlas - Test Writing42.6%Extra highGPT-5.5 (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 2 (Scale rank accounts for CI); ±5.96; entry added 2026-05-07
SWE-Bench Pro (public, v1)58.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—public set
SWE-Bench Pro (public, v1)59.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
SWE-bench Verified82.6%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.695 stderr; $1.362/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
SWE-bench Verified80.6%Extra highgpt-5.5-pre-release_xhighIndependent testIndependentEpoch AI ↗24 Apr 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv); pre-release checkpoint
tau2-bench91.8%MediumGPT-5.5 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated)
tau2-bench93.0%HighGPT-5.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated)
tau2-bench93.9%Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated)
Terminal-Bench 2.082.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026CodexTerminal-Bench 2.0
Terminal-Bench 2.180.5%MediumGPT-5.5 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated)
Terminal-Bench 2.179.4%HighGPT-5.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated)
Terminal-Bench 2.176.4%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.649 stderr; $0.742/test
Terminal-Bench 2.184.3%Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated)
Terminal-Bench 2.183.1%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Codex CLITB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.178.0%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Terminus 2TB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.185.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026CodexTerminal-Bench 2.1
Terminal-Bench 4.05.1%MediumGPT-5.5 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-medium) (AA marks this variant deprecated)
Terminal-Bench 4.09.1%HighGPT-5.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5-high) (AA marks this variant deprecated)
Terminal-Bench 4.014.6%Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-5) (AA marks this variant deprecated)
Toolathlon55.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—
Toolathlon55.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
Vals Code Migration45.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—±4.164 stderr; $6.442/test
Vals CorpFin v268.4%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗12 Aug 2026—vals id openai/gpt-5.5; rank 9/134; ±0.916 stderr; $0.43166/test
Vals SRE Bench3.8%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—±1.186 stderr; $17.774/test
Vectara Hallucination Leaderboard (HHEM)9.3%Defaultopenai/gpt-5.5Independent testIndependentVectara ↗22 Sep 2026—factual consistency 90.7 %; answer rate 100.0 %; avg summary 129.6 words; HHEM-2.3 judge; effort not stated (API default)
τ²-bench (Telecom)98.0%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗23 Apr 2026—original prompts, no prompt tuning, GPT-4.1 user model