Skip to content
Bencher

Models · OpenAI · Out sinceReleased 9 Jul 2026

GPT-5.6 Sol#9 for writing.#9 for writing, best at extra high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-5.6 Sol is made by OpenAI. Among the models we track it ranks #9 for writing, #11 for coding, #14 for research and analysis. It's expensive to use.

Launched at $5/$30; promotional $4/$20 since Aug 21, 2026 (through at least Nov 21, 2026) per OpenAI docs; OpenRouter lists $2/$10. Alias gpt-5.6. Default effort medium. "ultra" = multi-agent mode (4 parallel agents) in ChatGPT/Codex. Preview Jun 26, GA Jul 9, 2026.

Writing & creativity
61.3 / 100 · #9
Research & analysis
71.1 / 100 · #14
Coding
58.3 / 100 · #11
Price
Expensive$4 / $20
Price per 1M (blended)Blended / 1M
$8
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
217 (136 independent136 indep.)
Out sinceReleased
9 Jul 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowMediumHighExtra highMax

Highlighted: where it did best for writing (Extra high thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.6-sol.

Route it as openai/gpt-5.6-sol at $2 in / $10 out per 1M tokens, 1.1M context. Listed since 9 Jul 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-5.6 SolWhere GPT-5.6 Solthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Artificial Analysis Coding Agent Index

Best at High, worse abovePeaks at High

Terminal-Bench 4.0

Best at MaxBest at Max

CursorBench

Best at MaxBest at Max

DeepSWE

Best at MaxBest at Max

FrontierCode

Best at MaxBest at Max

SWE Atlas - Codebase QnA

Best at MaxBest at Max

ProgramBench (fully resolved)

Best at MaxBest at Max

Terminal-Bench 2.1

Best at Extra high, worse abovePeaks at Extra high

SciCode

Best at High, worse abovePeaks at High

Agents' Last Exam

Best at Extra high, worse abovePeaks at Extra high

Artificial Analysis Intelligence Index

Best at MaxBest at Max

AutomationBench

Best at MaxBest at Max

OSWorld 2.0 offline set

Best at MaxBest at Max

tau2-bench

Best at MaxBest at Max

ARC-AGI-2

Best at MaxBest at Max

ARC-AGI-3

Best at MaxBest at Max

GDPval-AA v2.1

Best at MaxBest at Max

GPQA Diamond

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

LLM Creative Story-Writing Benchmark (Lech Mazur)

Best at Extra highBest at Extra high

ScreenSpot-Pro

Best at Extra high, worse abovePeaks at Extra high

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent47.5%MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
Agents' Last Exam45.1%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.6324
Agents' Last Exam52.1%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $3.4126
Agents' Last Exam52.4%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $3.7874
Agents' Last Exam53.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $5.0764
Agents' Last Exam52.8%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $7.1322
Agents' Last Exam52.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
ARC-AGI-197.5%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
ARC-AGI-285.4%HighGPT-5.6 Sol (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.74/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-290.0%Extra highGPT-5.6 Sol (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.04/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-292.5%MaxGPT-5.6 Sol (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.44/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-292.5%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
ARC-AGI-30.3%LowGPT-5.6 Sol (Low)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $12,806; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-31.1%MediumGPT-5.6 Sol (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $12,971; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-32.1%HighGPT-5.6 Sol (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $15,176; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-37.0%Extra highGPT-5.6 Sol (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $19,216; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-37.8%MaxGPT-5.6 Sol (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $25,064; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-37.8%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
ARC-AGI-37.8%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
Artificial Analysis Coding Agent Index43.4No reasoningreasoning effort=noneMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $1.09
Artificial Analysis Coding Agent Index55.2Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $1.29
Artificial Analysis Coding Agent Index61.6Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $2.19
Artificial Analysis Coding Agent Index64.1Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3
Artificial Analysis Coding Agent Index63.3Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.74
Artificial Analysis Coding Agent Index54.6MaxGPT-5.6 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 72.3, SWE-Atlas-QnA 54.0, Terminal-Bench v4 37.4; avg cost $6.35/task; avg wall time 21 min/task
Artificial Analysis Coding Agent Index65.1Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $5
Artificial Analysis Coding Agent Index v1.180DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—Artificial Analysis Coding Agent Index v1.1 (index score)
Artificial Analysis Intelligence Index33.5LowGPT-5.6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-sol-low; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $0.26/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index39.2MediumGPT-5.6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-sol-medium; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $0.50/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index42.3HighGPT-5.6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-sol-high; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $0.81/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index44Extra highGPT-5.6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-sol-xhigh; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $1.18/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index47MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-sol; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $1.99/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index60.9Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—v4.1.1
Artificial Analysis Intelligence Index58.9DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—Artificial Analysis Intelligence Index v4.1
Artificial Analysis output speed67 tok/sLowGPT-5.6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 2.4s; list price $4/20 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed70 tok/sMediumGPT-5.6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 6.7s; list price $4/20 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed74 tok/sHighGPT-5.6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 16.4s; list price $4/20 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed76 tok/sExtra highGPT-5.6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 53.9s; list price $4/20 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed79 tok/sMaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 119.5s; list price $4/20 per 1M in/out (AA marks this variant deprecated)
AutomationBench11.7%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.31
AutomationBench19.6%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.42
AutomationBench24.8%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.47
AutomationBench26.3%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.54
AutomationBench28.8%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.67
AutomationBench18.1%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
BrowseComp92.2%DefaultGPT-5.6 Sol Ultra (4 parallel agents)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—ultra = multi-agent parallel test-time compute
BrowseComp90.4%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
BrowseComp90.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
CursorBench24.6%LowGPT-5.6 Sol LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 59; $0.87/task; 4,885 tokens/task; 21 steps/task
CursorBench24.6%Lowreasoning effort=lowIndependent testIndependentCursor (via Anthropic launch post) ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $0.87 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models
CursorBench31.1%MediumGPT-5.6 Sol MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 48; $1.77/task; 10,111 tokens/task; 32 steps/task
CursorBench31.1%Mediumreasoning effort=medIndependent testIndependentCursor (via Anthropic launch post) ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $1.77 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models
CursorBench35.7%HighGPT-5.6 Sol HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 38; $2.85/task; 16,174 tokens/task; 41 steps/task
CursorBench35.7%Highreasoning effort=highIndependent testIndependentCursor (via Anthropic launch post) ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $2.85 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models
CursorBench37.7%Extra highGPT-5.6 Sol Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 31; $4.40/task; 24,729 tokens/task; 55 steps/task
CursorBench37.7%Extra highreasoning effort=xhighIndependent testIndependentCursor (via Anthropic launch post) ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $4.4 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models
CursorBench41.7%MaxGPT-5.6 Sol MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 21; $8.23/task; 42,944 tokens/task; 99 steps/task
CursorBench41.7%Maxreasoning effort=maxIndependent testIndependentCursor (via Anthropic launch post) ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $8.23 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models
CursorBench 3.2.067.2%Maxreasoning effort=maxIndependent testIndependentCursor (via Anthropic launch post) ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; cost $5.69/task
DeepSWE45.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.82
DeepSWE61.1%Mediumgpt-5.6-sol_mediumIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 80.5%; ±1.6; 4 runs; $1.86/task
DeepSWE61.1%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.42
DeepSWE69.4%Highgpt-5.6-sol_highIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 86.7%; ±1.4; 4 runs; $3.47/task
DeepSWE69.4%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $2.66
DeepSWE70.7%Extra highgpt-5.6-sol_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.8%; ±0.8; 4 runs; $4.70/task
DeepSWE70.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.6
DeepSWE72.7%Maxgpt-5.6-sol_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.8%; ±2.8; 4 runs; $8.39/task
DeepSWE72.3%MaxGPT-5.6 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE72.7%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $6.46
DeepSWE72.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026Codex
Design Arena (all categories)1345Extra highgpt-5.6-sol-xhighIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±3.7 SE; 9719 battles; win rate 58.6%
Design Arena (all categories)1325Defaultgpt-5.6-solIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.5 SE; 24080 battles; win rate 57.4%
Design Arena (fullstack)1259Extra highgpt-5.6-sol-xhighIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±8.5 SE; 1973 battles; win rate 54.3%
Design Arena (fullstack)1157Defaultgpt-5.6-solIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±9.2 SE; 1614 battles; win rate 49.8%
Epoch Capabilities Index161.8DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 159.3-165.3; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1972Defaultgpt-5.6-solIndependent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 10; rubric score 16.78/20; slop 11.68; avg length 8548 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
EuroEval Swedish (generative)1.3Defaultgpt-5.6-sol (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 2; ±0.04; lower is better; task scores (first metric): SweDN summarisation 37.92 ± 0.22, Skolprov 90.97 ± 1.67, Swedish facts 85.67 ± 1.96, ScaLA-sv 72.02 ± 1.52
FrontierCode35.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $1.89
FrontierCode39.9%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.69
FrontierCode45.1%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $3.48
FrontierCode46.8%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $4.15
FrontierCode47.5%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $5.19
FrontierCode47.5%Defaultgpt-5.6-sol_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026codexread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5
FrontierCode v1.1 (Extended)60.6%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
FrontierMath (Tiers 1-3)89.1%Maxgpt-5.6-sol_maxIndependent testIndependentEpoch AI ↗9 Jul 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 1-389.0%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
FrontierMath Tier 482.9%Maxgpt-5.6-sol_maxIndependent testIndependentEpoch AI ↗9 Jul 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 5.9pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 483.0%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—v2
FrontierMath Tier 483.0%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
FrontierSWE v232.2%Maxreasoning effort=maxIndependent testIndependentProximal (via Anthropic system card) ↗22 Sep 2026Proximal harness
GDP.pdf30.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GDPval-AA v21711Defaultas listed by AnthropicIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗1 Sep 2026—
GDPval-AA v21748DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GDPval-AA v2.11289Lowreasoning effort=lowIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $0.27
GDPval-AA v2.11403Mediumreasoning effort=medIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $0.6
GDPval-AA v2.11480Highreasoning effort=highIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $1.11
GDPval-AA v2.11548Extra highreasoning effort=xhighIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $1.7
GDPval-AA v2.11588Maxreasoning effort=maxIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $2.81
GPQA Diamond89.8%LowGPT-5.6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated)
GPQA Diamond91.7%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.0197313
GPQA Diamond92.6%MediumGPT-5.6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated)
GPQA Diamond92.9%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.0298166
GPQA Diamond92.8%HighGPT-5.6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated)
GPQA Diamond93.2%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.048927
GPQA Diamond93.1%Extra highGPT-5.6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated)
GPQA Diamond93.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.0687297
GPQA Diamond95.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±1.074 stderr; $0.110/test
GPQA Diamond94.1%MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated)
GPQA Diamond93.5%Maxgpt-5.6-sol_maxIndependent testIndependentEpoch AI ↗9 Jul 2026—Epoch-run GPQA Diamond; stderr 1.6pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond94.6%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.1267
GPQA Diamond94.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GSO76.5%Extra highreasoning_effort=xhighIndependent testIndependentGSO ↗27 Sep 2026OpenHandsOpt@1; hack-controlled score 70.59; data https://gso-bench.github.io/assets/leaderboard.json
HealthBench Professional60.5%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—length-adjusted
HealthBench Professional60.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—official HealthBench Professional scoring
HLE Diamond31.2%Defaultgpt-5.6-solIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 7 (Scale rank accounts for CI); ±2.9; entry added 2025-11-19
Humanity's Last Exam39.4%LowGPT-5.6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated)
Humanity's Last Exam42.2%MediumGPT-5.6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated)
Humanity's Last Exam46.0%HighGPT-5.6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated)
Humanity's Last Exam47.3%Extra highGPT-5.6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated)
Humanity's Last Exam49.5%MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated)
IOI (Vals)91.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.51 stderr; $7.982/test
LegalBench (Vals)87.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-5.6-sol; rank 9/149; ±0.412 stderr; $0.012678/test
LLM Creative Story-Writing Benchmark (Lech Mazur)2.3HighGPT-5.6 Sol (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 11/56; Thurstone comparison score (centered at 0); est. win chance 81%; 95% bootstrap 2.217 to 2.368
LLM Creative Story-Writing Benchmark (Lech Mazur)2.5Extra highGPT-5.6 Sol (xhigh)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 8/56; Thurstone comparison score (centered at 0); est. win chance 83%; 95% bootstrap 2.407 to 2.561
LMArena Code Arena (WebDev)1619Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026CodexCode Arena | WebDev overall (agentic web-dev, raw); rank 22 (CI rank 15-24); 95% CI 1613-1625; 16911 votes
LMArena Search Arena1257Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗24 Aug 2026—rank 1 (rank range 1-2); 95% CI 1249.8-1264.7; 29663 votes
LMArena Text - Coding category1534Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 14 (CI rank 6-33); 95% CI 1527-1541; 9788 votes
LMArena Text - Creative Writing1469Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 7-35); 95% CI 1460.9-1477.0; 7388 votes; style-controlled
LMArena Text - Expert1535Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 1-23); 95% CI 1524.6-1544.4; 4053 votes; style-controlled
LMArena Text - Hard Prompts1511Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 7-32); 95% CI 1505.6-1516.0; 23609 votes; style-controlled
LMArena Text - Instruction Following1491Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 11 (rank range 6-22); 95% CI 1484.2-1497.0; 12619 votes; style-controlled
LMArena Text - Longer Query1498Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 7-31); 95% CI 1491.9-1503.9; 16784 votes; style-controlled
LMArena Text - Non-English1475Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 20 (rank range 8-34); 95% CI 1469.2-1479.8; 21430 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1485Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 20 (rank range 6-43); 95% CI 1476.5-1492.5; 6778 votes; style-controlled
LMArena Text - Occupational: Legal & Government1494Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 20 (rank range 2-57); 95% CI 1482.9-1505.7; 3029 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1482Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 6-28); 95% CI 1474.8-1489.1; 9675 votes; style-controlled
LMArena Text (overall)1484Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 20 (CI rank 11-34); 95% CI 1479-1488; 36002 votes
MCP Atlas81.8%Defaultgpt-5.6 (sol)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±2.4; entry added 2026-07-15
MMMU-Pro (no tools)83.0%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
MMMU-Pro (with tools)84.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OpenAI MRCR v2 (8-needle)91.5%Defaultbest score across efforts (OpenAI table: "maximum at any effort"); 8-needle 256K-512KMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
OpenAI MRCR v2 (8-needle)91.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 256K-512KMaker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OpenAI MRCR v2 (8-needle)73.8%Defaultbest score across efforts (OpenAI table: "maximum at any effort"); 8-needle 512K-1MMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
OpenAI MRCR v2 (8-needle)73.8%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 512K-1MMaker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OSWorld 2.027.3%Maxgpt-5.6-sol_maxIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6272; tool setting batch tool; step budget 500
OSWorld 2.062.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—OSWorld 2.0 (OpenAI-run)
OSWorld 2.0 offline set16.4%No reasoningreasoning effort=noneMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.78171
OSWorld 2.0 offline set29.8%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.9086
OSWorld 2.0 offline set49.7%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $2.7275
OSWorld 2.0 offline set56.5%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $4.4613
OSWorld 2.0 offline set60.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $5.9257
OSWorld 2.0 offline set66.2%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $7.7105
PRBench Finance (Scale)50.5%Maxgpt-5.6-sol (max)Independent testIndependentScale AI (SEAL) ↗28 Jul 2026—Scale rank 7; ±0.33 CI
PRBench Legal (Scale)50.5%Maxgpt-5.6-sol (max)Independent testIndependentScale AI (SEAL) ↗28 Jul 2026—Scale rank 7; ±0.54 CI
ProgramBench (avg test pass rate)69.9%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 1.0%, almost (>=95% tests) 15.5%; avg cost $6.08/task
ProgramBench (avg test pass rate)57.8%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.5%, almost (>=95% tests) 2.5%; avg cost $1.00/task
ProgramBench (fully resolved)1.0%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 15.5%; avg cost $6.08/task
ProgramBench (fully resolved)1.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.862 stderr; $15.289/test; strict fully-resolved rate
ProgramBench (fully resolved)0.5%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 2.5%; avg cost $1.00/task
SciCode56.4%LowGPT-5.6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated)
SciCode57.4%MediumGPT-5.6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated)
SciCode57.8%HighGPT-5.6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated)
SciCode57.1%Extra highGPT-5.6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated)
SciCode57.1%MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated)
ScreenSpot-Pro59.3%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.03213
ScreenSpot-Pro68.7%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.03398
ScreenSpot-Pro75.1%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.03648
ScreenSpot-Pro76.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.04098
ScreenSpot-Pro76.8%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.04931
SimpleBench64.8%Extra highGPT-5.6 Sol (xhigh)Independent testIndependentSimpleBench ↗9 Jul 2026—AVG@5, temp 0.7; rank 24th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified69.7%Maxgpt-5.6-sol_maxIndependent testIndependentEpoch AI ↗10 Aug 2026—Epoch-run (no tools); ±1.45 stderr
SkillsBench54.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.73 stderr; $4.140/test
SWE Atlas - Codebase QnA46.0%Extra highGPT-5.6-Sol (Codex) xHigh*Independent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 2 (Scale rank accounts for CI); ±5; entry added 2026-03-02; * = Scale footnote (see page)
SWE Atlas - Codebase QnA54.0%MaxGPT-5.6 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Test Writing45.9%Extra highGPT-5.6-Sol (Codex) xHigh*Independent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 1 (Scale rank accounts for CI); ±6; entry added 2026-07-28; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)64.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
SWE-Bench Pro V2 (full)95.5%Extra highGPT-5.6-sol (Codex) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 6 (Scale rank accounts for CI); ±1.4; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)82.4%Extra highGPT-5.6 Sol (Codex) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 8 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-bench Verified96.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±0.856 stderr; $1.151/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
SWE-rebench62.3%MediumGPT-5.6 Sol [medium]Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.83; pass@5 79.3%; $0.85/problem
tau2-bench76.0%LowGPT-5.6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated)
tau2-bench81.0%MediumGPT-5.6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated)
tau2-bench83.3%HighGPT-5.6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated)
tau2-bench84.8%Extra highGPT-5.6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated)
tau2-bench85.1%MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated)
Terminal-Bench 2.176.8%LowGPT-5.6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated)
Terminal-Bench 2.186.1%MediumGPT-5.6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated)
Terminal-Bench 2.187.3%HighGPT-5.6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated)
Terminal-Bench 2.189.5%Extra highGPT-5.6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated)
Terminal-Bench 2.188.0%MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated)
Terminal-Bench 2.185.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.35 stderr; $1.023/test
Terminal-Bench 2.191.9%DefaultGPT-5.6 Sol Ultra (4 parallel agents)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026Codexultra = multi-agent parallel test-time compute (4 agents by default)
Terminal-Bench 2.188.8%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026CodexTerminal-Bench 2.1
Terminal-Bench 3.034.6%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026CodexTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 5.8B, run cost $4.0k
Terminal-Bench 4.01.0%LowGPT-5.6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated)
Terminal-Bench 4.07.9%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $1.46089
Terminal-Bench 4.014.6%MediumGPT-5.6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated)
Terminal-Bench 4.020.9%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $2.69056
Terminal-Bench 4.020.7%HighGPT-5.6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated)
Terminal-Bench 4.026.1%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $4.12252
Terminal-Bench 4.024.7%Extra highGPT-5.6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated)
Terminal-Bench 4.028.5%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $5.38518
Terminal-Bench 4.039.9%MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated)
Terminal-Bench 4.037.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0 stderr; $7.984/test
Terminal-Bench 4.037.4%MaxGPT-5.6 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.037.3%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026CodexTB 4.0.0 (66 tasks); ±3.78 95% CI; 330 trials; total run cost $2542; model release 2026-06-26
Terminal-Bench 4.037.3%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $7.89451
Terminal-Bench-Science 0.122.4%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $20.10413
Toolathlon58.0%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
Vals Code Migration52.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.353 stderr; $24.541/test
Vals Index58.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.017 stderr; $14.239/test
Vals Legal Research Bench48.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-5.6-sol; rank 9/72; ±3.473 stderr; $21.605566/test
Vals Public Benefits Bench v1.166.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-5.6-sol; rank 16/45; ±1.228 stderr; $8.847882/test
Vals SRE Bench30.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±2.851 stderr; $42.521/test
Vals Vibe Code Bench80.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±3.716 stderr; $33.404/test
Vectara Hallucination Leaderboard (HHEM)12.4%Defaultopenai/gpt-5.6-solIndependent testIndependentVectara ↗22 Sep 2026—factual consistency 87.6 %; answer rate 99.1 %; avg summary 131.1 words; HHEM-2.3 judge; effort not stated (API default)
Vending-Bench 2$9,619DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$1,338; rank 7; only top 10 rendered server-side (57 more behind 'Show more')