Models · OpenAI · Out sinceReleased 9 Jul 2026
GPT-5.6 Sol#9 for writing.#9 for writing, best at extra high effort.
GPT-5.6 Sol is made by OpenAI. Among the models we track it ranks #9 for writing, #11 for coding, #14 for research and analysis. It's expensive to use.
Launched at $5/$30; promotional $4/$20 since Aug 21, 2026 (through at least Nov 21, 2026) per OpenAI docs; OpenRouter lists $2/$10. Alias gpt-5.6. Default effort medium. "ultra" = multi-agent mode (4 parallel agents) in ChatGPT/Codex. Preview Jun 26, GA Jul 9, 2026.
- 61.3 / 100 · #9
- 71.1 / 100 · #14
- 58.3 / 100 · #11
- Expensive$4 / $20
- $8
- 1.1M
- 128K
- 217 (136 independent136 indep.)
- 9 Jul 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (Extra high thinking).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.6-sol.
Route it as openai/gpt-5.6-sol at $2 in / $10 out per 1M tokens, 1.1M context. Listed since 9 Jul 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-5.6 SolWhere GPT-5.6 Solthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Artificial Analysis Coding Agent Index
Best at High, worse abovePeaks at High
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
DeepSWE
Best at MaxBest at Max
FrontierCode
Best at MaxBest at Max
SWE Atlas - Codebase QnA
Best at MaxBest at Max
ProgramBench (fully resolved)
Best at MaxBest at Max
Terminal-Bench 2.1
Best at Extra high, worse abovePeaks at Extra high
SciCode
Best at High, worse abovePeaks at High
Agents' Last Exam
Best at Extra high, worse abovePeaks at Extra high
Artificial Analysis Intelligence Index
Best at MaxBest at Max
AutomationBench
Best at MaxBest at Max
OSWorld 2.0 offline set
Best at MaxBest at Max
tau2-bench
Best at MaxBest at Max
ARC-AGI-2
Best at MaxBest at Max
ARC-AGI-3
Best at MaxBest at Max
GDPval-AA v2.1
Best at MaxBest at Max
GPQA Diamond
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
LLM Creative Story-Writing Benchmark (Lech Mazur)
Best at Extra highBest at Extra high
ScreenSpot-Pro
Best at Extra high, worse abovePeaks at Extra high
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 47.5% | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| Agents' Last Exam | 45.1% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.6324 |
| Agents' Last Exam | 52.1% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $3.4126 |
| Agents' Last Exam | 52.4% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $3.7874 |
| Agents' Last Exam | 53.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $5.0764 |
| Agents' Last Exam | 52.8% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $7.1322 |
| Agents' Last Exam | 52.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| ARC-AGI-1 | 97.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| ARC-AGI-2 | 85.4% | HighGPT-5.6 Sol (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.74/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 90.0% | Extra highGPT-5.6 Sol (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.04/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 92.5% | MaxGPT-5.6 Sol (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.44/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 92.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| ARC-AGI-3 | 0.3% | LowGPT-5.6 Sol (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $12,806; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 1.1% | MediumGPT-5.6 Sol (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $12,971; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 2.1% | HighGPT-5.6 Sol (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $15,176; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 7.0% | Extra highGPT-5.6 Sol (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $19,216; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 7.8% | MaxGPT-5.6 Sol (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $25,064; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 7.8% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| ARC-AGI-3 | 7.8% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| Artificial Analysis Coding Agent Index | 43.4 | No reasoningreasoning effort=none | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $1.09 |
| Artificial Analysis Coding Agent Index | 55.2 | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $1.29 |
| Artificial Analysis Coding Agent Index | 61.6 | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $2.19 |
| Artificial Analysis Coding Agent Index | 64.1 | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3 |
| Artificial Analysis Coding Agent Index | 63.3 | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.74 |
| Artificial Analysis Coding Agent Index | 54.6 | MaxGPT-5.6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 72.3, SWE-Atlas-QnA 54.0, Terminal-Bench v4 37.4; avg cost $6.35/task; avg wall time 21 min/task |
| Artificial Analysis Coding Agent Index | 65.1 | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $5 |
| Artificial Analysis Coding Agent Index v1.1 | 80 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | Artificial Analysis Coding Agent Index v1.1 (index score) |
| Artificial Analysis Intelligence Index | 33.5 | LowGPT-5.6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-sol-low; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $0.26/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 39.2 | MediumGPT-5.6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-sol-medium; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $0.50/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 42.3 | HighGPT-5.6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-sol-high; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $0.81/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 44 | Extra highGPT-5.6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-sol-xhigh; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $1.18/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 47 | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-sol; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $1.99/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 60.9 | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | v4.1.1 |
| Artificial Analysis Intelligence Index | 58.9 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | Artificial Analysis Intelligence Index v4.1 |
| Artificial Analysis output speed | 67 tok/s | LowGPT-5.6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 2.4s; list price $4/20 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 70 tok/s | MediumGPT-5.6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 6.7s; list price $4/20 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 74 tok/s | HighGPT-5.6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 16.4s; list price $4/20 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 76 tok/s | Extra highGPT-5.6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 53.9s; list price $4/20 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 79 tok/s | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 119.5s; list price $4/20 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 11.7% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.31 |
| AutomationBench | 19.6% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.42 |
| AutomationBench | 24.8% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.47 |
| AutomationBench | 26.3% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.54 |
| AutomationBench | 28.8% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.67 |
| AutomationBench | 18.1% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| BrowseComp | 92.2% | DefaultGPT-5.6 Sol Ultra (4 parallel agents) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | ultra = multi-agent parallel test-time compute |
| BrowseComp | 90.4% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| BrowseComp | 90.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| CursorBench | 24.6% | LowGPT-5.6 Sol Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 59; $0.87/task; 4,885 tokens/task; 21 steps/task |
| CursorBench | 24.6% | Lowreasoning effort=low | Independent testIndependentCursor (via Anthropic launch post) ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $0.87 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models |
| CursorBench | 31.1% | MediumGPT-5.6 Sol Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 48; $1.77/task; 10,111 tokens/task; 32 steps/task |
| CursorBench | 31.1% | Mediumreasoning effort=med | Independent testIndependentCursor (via Anthropic launch post) ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $1.77 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models |
| CursorBench | 35.7% | HighGPT-5.6 Sol High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 38; $2.85/task; 16,174 tokens/task; 41 steps/task |
| CursorBench | 35.7% | Highreasoning effort=high | Independent testIndependentCursor (via Anthropic launch post) ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $2.85 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models |
| CursorBench | 37.7% | Extra highGPT-5.6 Sol Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 31; $4.40/task; 24,729 tokens/task; 55 steps/task |
| CursorBench | 37.7% | Extra highreasoning effort=xhigh | Independent testIndependentCursor (via Anthropic launch post) ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $4.4 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models |
| CursorBench | 41.7% | MaxGPT-5.6 Sol Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 21; $8.23/task; 42,944 tokens/task; 99 steps/task |
| CursorBench | 41.7% | Maxreasoning effort=max | Independent testIndependentCursor (via Anthropic launch post) ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; cost/task $8.23 as published by Cursor; OpenAI has not published CursorBench 4.0 for GPT-6 models |
| CursorBench 3.2.0 | 67.2% | Maxreasoning effort=max | Independent testIndependentCursor (via Anthropic launch post) ↗ | 1 Sep 2026 | Cursor agent | CursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; cost $5.69/task |
| DeepSWE | 45.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.82 |
| DeepSWE | 61.1% | Mediumgpt-5.6-sol_medium | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 80.5%; ±1.6; 4 runs; $1.86/task |
| DeepSWE | 61.1% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.42 |
| DeepSWE | 69.4% | Highgpt-5.6-sol_high | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 86.7%; ±1.4; 4 runs; $3.47/task |
| DeepSWE | 69.4% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $2.66 |
| DeepSWE | 70.7% | Extra highgpt-5.6-sol_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.8%; ±0.8; 4 runs; $4.70/task |
| DeepSWE | 70.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.6 |
| DeepSWE | 72.7% | Maxgpt-5.6-sol_max | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.8%; ±2.8; 4 runs; $8.39/task |
| DeepSWE | 72.3% | MaxGPT-5.6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 72.7% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $6.46 |
| DeepSWE | 72.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | |
| Design Arena (all categories) | 1345 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±3.7 SE; 9719 battles; win rate 58.6% |
| Design Arena (all categories) | 1325 | Defaultgpt-5.6-sol | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.5 SE; 24080 battles; win rate 57.4% |
| Design Arena (fullstack) | 1259 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±8.5 SE; 1973 battles; win rate 54.3% |
| Design Arena (fullstack) | 1157 | Defaultgpt-5.6-sol | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±9.2 SE; 1614 battles; win rate 49.8% |
| Epoch Capabilities Index | 161.8 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 159.3-165.3; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1972 | Defaultgpt-5.6-sol | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 10; rubric score 16.78/20; slop 11.68; avg length 8548 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| EuroEval Swedish (generative) | 1.3 | Defaultgpt-5.6-sol (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 2; ±0.04; lower is better; task scores (first metric): SweDN summarisation 37.92 ± 0.22, Skolprov 90.97 ± 1.67, Swedish facts 85.67 ± 1.96, ScaLA-sv 72.02 ± 1.52 |
| FrontierCode | 35.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $1.89 |
| FrontierCode | 39.9% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.69 |
| FrontierCode | 45.1% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $3.48 |
| FrontierCode | 46.8% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $4.15 |
| FrontierCode | 47.5% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $5.19 |
| FrontierCode | 47.5% | Defaultgpt-5.6-sol_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | codex | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5 |
| FrontierCode v1.1 (Extended) | 60.6% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| FrontierMath (Tiers 1-3) | 89.1% | Maxgpt-5.6-sol_max | Independent testIndependentEpoch AI ↗ | 9 Jul 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 1-3 | 89.0% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| FrontierMath Tier 4 | 82.9% | Maxgpt-5.6-sol_max | Independent testIndependentEpoch AI ↗ | 9 Jul 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 5.9pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 83.0% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | v2 |
| FrontierMath Tier 4 | 83.0% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| FrontierSWE v2 | 32.2% | Maxreasoning effort=max | Independent testIndependentProximal (via Anthropic system card) ↗ | 22 Sep 2026 | Proximal harness | |
| GDP.pdf | 30.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GDPval-AA v2 | 1711 | Defaultas listed by Anthropic | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 1 Sep 2026 | — | |
| GDPval-AA v2 | 1748 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GDPval-AA v2.1 | 1289 | Lowreasoning effort=low | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $0.27 |
| GDPval-AA v2.1 | 1403 | Mediumreasoning effort=med | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $0.6 |
| GDPval-AA v2.1 | 1480 | Highreasoning effort=high | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $1.11 |
| GDPval-AA v2.1 | 1548 | Extra highreasoning effort=xhigh | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $1.7 |
| GDPval-AA v2.1 | 1588 | Maxreasoning effort=max | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $2.81 |
| GPQA Diamond | 89.8% | LowGPT-5.6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated) |
| GPQA Diamond | 91.7% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.0197313 |
| GPQA Diamond | 92.6% | MediumGPT-5.6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated) |
| GPQA Diamond | 92.9% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.0298166 |
| GPQA Diamond | 92.8% | HighGPT-5.6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated) |
| GPQA Diamond | 93.2% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.048927 |
| GPQA Diamond | 93.1% | Extra highGPT-5.6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated) |
| GPQA Diamond | 93.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.0687297 |
| GPQA Diamond | 95.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.074 stderr; $0.110/test |
| GPQA Diamond | 94.1% | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated) |
| GPQA Diamond | 93.5% | Maxgpt-5.6-sol_max | Independent testIndependentEpoch AI ↗ | 9 Jul 2026 | — | Epoch-run GPQA Diamond; stderr 1.6pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 94.6% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.1267 |
| GPQA Diamond | 94.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GSO | 76.5% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | 27 Sep 2026 | OpenHands | Opt@1; hack-controlled score 70.59; data https://gso-bench.github.io/assets/leaderboard.json |
| HealthBench Professional | 60.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | length-adjusted |
| HealthBench Professional | 60.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | official HealthBench Professional scoring |
| HLE Diamond | 31.2% | Defaultgpt-5.6-sol | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 7 (Scale rank accounts for CI); ±2.9; entry added 2025-11-19 |
| Humanity's Last Exam | 39.4% | LowGPT-5.6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated) |
| Humanity's Last Exam | 42.2% | MediumGPT-5.6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated) |
| Humanity's Last Exam | 46.0% | HighGPT-5.6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated) |
| Humanity's Last Exam | 47.3% | Extra highGPT-5.6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated) |
| Humanity's Last Exam | 49.5% | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated) |
| IOI (Vals) | 91.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.51 stderr; $7.982/test |
| LegalBench (Vals) | 87.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-5.6-sol; rank 9/149; ±0.412 stderr; $0.012678/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 2.3 | HighGPT-5.6 Sol (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 11/56; Thurstone comparison score (centered at 0); est. win chance 81%; 95% bootstrap 2.217 to 2.368 |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 2.5 | Extra highGPT-5.6 Sol (xhigh) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 8/56; Thurstone comparison score (centered at 0); est. win chance 83%; 95% bootstrap 2.407 to 2.561 |
| LMArena Code Arena (WebDev) | 1619 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | Codex | Code Arena | WebDev overall (agentic web-dev, raw); rank 22 (CI rank 15-24); 95% CI 1613-1625; 16911 votes |
| LMArena Search Arena | 1257 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 1 (rank range 1-2); 95% CI 1249.8-1264.7; 29663 votes |
| LMArena Text - Coding category | 1534 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 14 (CI rank 6-33); 95% CI 1527-1541; 9788 votes |
| LMArena Text - Creative Writing | 1469 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 19 (rank range 7-35); 95% CI 1460.9-1477.0; 7388 votes; style-controlled |
| LMArena Text - Expert | 1535 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 6 (rank range 1-23); 95% CI 1524.6-1544.4; 4053 votes; style-controlled |
| LMArena Text - Hard Prompts | 1511 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 16 (rank range 7-32); 95% CI 1505.6-1516.0; 23609 votes; style-controlled |
| LMArena Text - Instruction Following | 1491 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 11 (rank range 6-22); 95% CI 1484.2-1497.0; 12619 votes; style-controlled |
| LMArena Text - Longer Query | 1498 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 7-31); 95% CI 1491.9-1503.9; 16784 votes; style-controlled |
| LMArena Text - Non-English | 1475 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 20 (rank range 8-34); 95% CI 1469.2-1479.8; 21430 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1485 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 20 (rank range 6-43); 95% CI 1476.5-1492.5; 6778 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1494 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 20 (rank range 2-57); 95% CI 1482.9-1505.7; 3029 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1482 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 6-28); 95% CI 1474.8-1489.1; 9675 votes; style-controlled |
| LMArena Text (overall) | 1484 | Extra highgpt-5.6-sol-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 20 (CI rank 11-34); 95% CI 1479-1488; 36002 votes |
| MCP Atlas | 81.8% | Defaultgpt-5.6 (sol) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±2.4; entry added 2026-07-15 |
| MMMU-Pro (no tools) | 83.0% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| MMMU-Pro (with tools) | 84.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 91.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort"); 8-needle 256K-512K | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 91.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 256K-512K | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 73.8% | Defaultbest score across efforts (OpenAI table: "maximum at any effort"); 8-needle 512K-1M | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 73.8% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 512K-1M | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OSWorld 2.0 | 27.3% | Maxgpt-5.6-sol_max | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6272; tool setting batch tool; step budget 500 |
| OSWorld 2.0 | 62.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | OSWorld 2.0 (OpenAI-run) |
| OSWorld 2.0 offline set | 16.4% | No reasoningreasoning effort=none | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.78171 |
| OSWorld 2.0 offline set | 29.8% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.9086 |
| OSWorld 2.0 offline set | 49.7% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $2.7275 |
| OSWorld 2.0 offline set | 56.5% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $4.4613 |
| OSWorld 2.0 offline set | 60.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $5.9257 |
| OSWorld 2.0 offline set | 66.2% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $7.7105 |
| PRBench Finance (Scale) | 50.5% | Maxgpt-5.6-sol (max) | Independent testIndependentScale AI (SEAL) ↗ | 28 Jul 2026 | — | Scale rank 7; ±0.33 CI |
| PRBench Legal (Scale) | 50.5% | Maxgpt-5.6-sol (max) | Independent testIndependentScale AI (SEAL) ↗ | 28 Jul 2026 | — | Scale rank 7; ±0.54 CI |
| ProgramBench (avg test pass rate) | 69.9% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 1.0%, almost (>=95% tests) 15.5%; avg cost $6.08/task |
| ProgramBench (avg test pass rate) | 57.8% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.5%, almost (>=95% tests) 2.5%; avg cost $1.00/task |
| ProgramBench (fully resolved) | 1.0% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 15.5%; avg cost $6.08/task |
| ProgramBench (fully resolved) | 1.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.862 stderr; $15.289/test; strict fully-resolved rate |
| ProgramBench (fully resolved) | 0.5% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 2.5%; avg cost $1.00/task |
| SciCode | 56.4% | LowGPT-5.6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated) |
| SciCode | 57.4% | MediumGPT-5.6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated) |
| SciCode | 57.8% | HighGPT-5.6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated) |
| SciCode | 57.1% | Extra highGPT-5.6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated) |
| SciCode | 57.1% | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated) |
| ScreenSpot-Pro | 59.3% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.03213 |
| ScreenSpot-Pro | 68.7% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.03398 |
| ScreenSpot-Pro | 75.1% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.03648 |
| ScreenSpot-Pro | 76.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.04098 |
| ScreenSpot-Pro | 76.8% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.04931 |
| SimpleBench | 64.8% | Extra highGPT-5.6 Sol (xhigh) | Independent testIndependentSimpleBench ↗ | 9 Jul 2026 | — | AVG@5, temp 0.7; rank 24th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 69.7% | Maxgpt-5.6-sol_max | Independent testIndependentEpoch AI ↗ | 10 Aug 2026 | — | Epoch-run (no tools); ±1.45 stderr |
| SkillsBench | 54.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.73 stderr; $4.140/test |
| SWE Atlas - Codebase QnA | 46.0% | Extra highGPT-5.6-Sol (Codex) xHigh* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 2 (Scale rank accounts for CI); ±5; entry added 2026-03-02; * = Scale footnote (see page) |
| SWE Atlas - Codebase QnA | 54.0% | MaxGPT-5.6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Test Writing | 45.9% | Extra highGPT-5.6-Sol (Codex) xHigh* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 1 (Scale rank accounts for CI); ±6; entry added 2026-07-28; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 64.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| SWE-Bench Pro V2 (full) | 95.5% | Extra highGPT-5.6-sol (Codex) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 6 (Scale rank accounts for CI); ±1.4; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-Bench Pro V2 (hard) | 82.4% | Extra highGPT-5.6 Sol (Codex) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 8 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-bench Verified | 96.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±0.856 stderr; $1.151/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| SWE-rebench | 62.3% | MediumGPT-5.6 Sol [medium] | Independent testIndependentSWE-rebench (Nebius) ↗ | 1 Oct 2026 | SWE-rebench standard scaffold | time window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.83; pass@5 79.3%; $0.85/problem |
| tau2-bench | 76.0% | LowGPT-5.6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated) |
| tau2-bench | 81.0% | MediumGPT-5.6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated) |
| tau2-bench | 83.3% | HighGPT-5.6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated) |
| tau2-bench | 84.8% | Extra highGPT-5.6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated) |
| tau2-bench | 85.1% | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 76.8% | LowGPT-5.6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 86.1% | MediumGPT-5.6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 87.3% | HighGPT-5.6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 89.5% | Extra highGPT-5.6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 88.0% | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 85.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.35 stderr; $1.023/test |
| Terminal-Bench 2.1 | 91.9% | DefaultGPT-5.6 Sol Ultra (4 parallel agents) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | ultra = multi-agent parallel test-time compute (4 agents by default) |
| Terminal-Bench 2.1 | 88.8% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | Terminal-Bench 2.1 |
| Terminal-Bench 3.0 | 34.6% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Codex | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 5.8B, run cost $4.0k |
| Terminal-Bench 4.0 | 1.0% | LowGPT-5.6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-low) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 7.9% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $1.46089 |
| Terminal-Bench 4.0 | 14.6% | MediumGPT-5.6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-medium) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 20.9% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $2.69056 |
| Terminal-Bench 4.0 | 20.7% | HighGPT-5.6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-high) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 26.1% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $4.12252 |
| Terminal-Bench 4.0 | 24.7% | Extra highGPT-5.6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol-xhigh) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 28.5% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $5.38518 |
| Terminal-Bench 4.0 | 39.9% | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-sol) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 37.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0 stderr; $7.984/test |
| Terminal-Bench 4.0 | 37.4% | MaxGPT-5.6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 37.3% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Codex | TB 4.0.0 (66 tasks); ±3.78 95% CI; 330 trials; total run cost $2542; model release 2026-06-26 |
| Terminal-Bench 4.0 | 37.3% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $7.89451 |
| Terminal-Bench-Science 0.1 | 22.4% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $20.10413 |
| Toolathlon | 58.0% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| Vals Code Migration | 52.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.353 stderr; $24.541/test |
| Vals Index | 58.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.017 stderr; $14.239/test |
| Vals Legal Research Bench | 48.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-5.6-sol; rank 9/72; ±3.473 stderr; $21.605566/test |
| Vals Public Benefits Bench v1.1 | 66.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-5.6-sol; rank 16/45; ±1.228 stderr; $8.847882/test |
| Vals SRE Bench | 30.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.851 stderr; $42.521/test |
| Vals Vibe Code Bench | 80.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±3.716 stderr; $33.404/test |
| Vectara Hallucination Leaderboard (HHEM) | 12.4% | Defaultopenai/gpt-5.6-sol | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 87.6 %; answer rate 99.1 %; avg summary 131.1 words; HHEM-2.3 judge; effort not stated (API default) |
| Vending-Bench 2 | $9,619 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$1,338; rank 7; only top 10 rendered server-side (57 more behind 'Show more') |