Models · OpenAI · Out sinceReleased 22 Sep 2026
GPT-6 Sol#10 for coding.#10 for coding, best at max effort.
GPT-6 Sol is made by OpenAI. Among the models we track it ranks #10 for coding, #28 for writing, #31 for research and analysis. It's mid-priced to use.
Superseded by GPT-6.1 Sol a week later. Default effort medium. Cached input $0.20/M. Knowledge cutoff Apr 20, 2026. System card link on launch post points to the GPT-6 Astra system card.
- 36.5 / 100 · #28
- 49.7 / 100 · #31
- 64.0 / 100 · #10
- Mid-priced$2 / $10
- $4
- 1.1M
- 128K
- 138 (103 independent103 indep.)
- 22 Sep 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-6-sol.
Route it as openai/gpt-6-sol at $2 in / $10 out per 1M tokens, 1.1M context. Listed since 22 Sep 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-6 SolWhere GPT-6 Solthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
DeepSWE
Best at MaxBest at Max
FrontierCode
Best at MaxBest at Max
SciCode
Best at MaxBest at Max
Agents' Last Exam
Best at MaxBest at Max
Artificial Analysis Intelligence Index
Best at MaxBest at Max
Terminal-Bench-Science 0.1
Best at MaxBest at Max
AutomationBench
Best at Extra high, worse abovePeaks at Extra high
OSWorld 2.0 offline set
Best at MaxBest at Max
AA-Briefcase v1.1
Best at MaxBest at Max
AA-LCR
Best at HighBest at High
AA-Omniscience Accuracy
Best at MaxBest at Max
AA-Omniscience Hallucination Rate
Best at Low, worse abovePeaks at Low
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
ARC-AGI-3
Best at MaxBest at Max
GDP.pdf
Best at High, worse abovePeaks at High
Humanity's Last Exam
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 905 | Lowreasoning effort=low | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $0.12 |
| AA-Briefcase v1.1 | 1142 | Mediumreasoning effort=med | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $0.34 |
| AA-Briefcase v1.1 | 1289 | Highreasoning effort=high | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $0.63 |
| AA-Briefcase v1.1 | 1364 | Extra highreasoning effort=xhigh | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $1.19 |
| AA-Briefcase v1.1 | 1483 | Maxreasoning effort=max | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $2.67 |
| AA-Briefcase v1.1 | 1479 | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 79.3% | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 82.3% | MediumGPT-6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.7% | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 81.3% | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.7% | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 51.3% | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 53.5% | MediumGPT-6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 53.7% | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 53.9% | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 54.5% | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 50.7% | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 56.8% | MediumGPT-6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 58.1% | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 58.9% | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 60.1% | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 26.5 | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 27 | MediumGPT-6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 26.8 | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 26.7 | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 27.1 | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 48.7% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.8641 |
| Agents' Last Exam | 53.1% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.2664 |
| Agents' Last Exam | 52.6% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.5302 |
| Agents' Last Exam | 55.4% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.668 |
| Agents' Last Exam | 56.4% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $2.9313 |
| ARC-AGI-2 | 89.6% | MaxGPT-6 Sol (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.44/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-3 | 1.2% | No reasoningGPT-6 Sol - Provider Adapter (None) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $4,234; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.3% | No reasoningGPT-6 Sol (None) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $4,656; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 1.8% | LowGPT-6 Sol - Provider Adapter (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $4,745; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 4.9% | MediumGPT-6 Sol - Provider Adapter (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $5,889; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.4% | MediumGPT-6 Sol (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $4,864; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 5.9% | HighGPT-6 Sol - Provider Adapter (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $5,908; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.9% | HighGPT-6 Sol (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $5,072; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 9.5% | Extra highGPT-6 Sol - Provider Adapter (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $6,907; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 1.8% | Extra highGPT-6 Sol (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $5,106; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 23.0% | MaxGPT-6 Sol - Provider Adapter (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $8,722; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 4.6% | MaxGPT-6 Sol (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $5,554; data https://arcprize.org/media/data/leaderboard/v3.json |
| Artificial Analysis Coding Agent Index | 56.7 | MaxGPT-6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 69.0, SWE-Atlas-QnA 57.5, Terminal-Bench v4 43.4; avg cost $2.99/task; avg wall time 22 min/task |
| Artificial Analysis Intelligence Index | 34.2 | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-sol-low; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.13/task |
| Artificial Analysis Intelligence Index | 39.8 | MediumGPT-6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-sol-medium; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.25/task |
| Artificial Analysis Intelligence Index | 42.4 | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-sol-high; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.37/task |
| Artificial Analysis Intelligence Index | 44.2 | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-sol-xhigh; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.52/task |
| Artificial Analysis Intelligence Index | 47.6 | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-sol; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $1.04/task |
| Artificial Analysis output speed | 56 tok/s | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 1.9s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 63 tok/s | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 19.2s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 68 tok/s | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 57.8s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 74 tok/s | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 177.8s; list price $2/10 per 1M in/out |
| AutomationBench | 21.2% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.1861 |
| AutomationBench | 26.9% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2098 |
| AutomationBench | 31.2% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2368 |
| AutomationBench | 33.2% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2746 |
| AutomationBench | 32.0% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.3406 |
| Chartography (no tools) | 53.6% | Defaultas listed by Anthropic (effort not stated) | Independent testIndependentSurge AI (via Anthropic launch post) ↗ | 28 Sep 2026 | — | no tools |
| DeepSWE | 37.2% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.1623 |
| DeepSWE | 56.6% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.3798 |
| DeepSWE | 65.3% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.6404 |
| DeepSWE | 66.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.0033 |
| DeepSWE | 69.0% | MaxGPT-6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 68.8% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $2.7439 |
| Design Arena (all categories) | 1292 | Defaultgpt-6-sol | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±7.8 SE; 2044 battles; win rate 52% |
| Design Arena (fullstack) | 1237 | Defaultgpt-6-sol | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±10.8 SE; 1154 battles; win rate 52.2% |
| EQ-Bench Creative Writing v3 (Elo) | 2125 | Default*gpt-6-sol | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 4; rubric score 16.51/20; slop 8.97; avg length 6182 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*) |
| FrontierCode | 37.3% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.453 |
| FrontierCode | 45.9% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.7966 |
| FrontierCode | 47.7% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $1.0793 |
| FrontierCode | 48.5% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $1.3748 |
| FrontierCode | 49.3% | Maxgpt-6-sol_max | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | codex | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5 |
| FrontierCode | 49.3% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.137 |
| FrontierMath (Tiers 1-3) | 89.8% | Maxgpt-6-sol_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 90.0% | Maxgpt-6-sol_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 5.2pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| GDP.pdf | 21.8% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.33 |
| GDP.pdf | 25.4% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.34 |
| GDP.pdf | 28.0% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.35 |
| GDP.pdf | 23.8% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.37 |
| GDP.pdf | 24.8% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.43 |
| GDPval-AA v2.1 | 1505 | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1480.19-1530.25 |
| GDPval-AA v2.1 | 1487 | Defaultas listed by Anthropic (effort not stated) | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 28 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); may predate OpenAI image-understanding bug fix for GPT-6 Sol |
| GPQA Diamond | 94.3% | Maxgpt-6-sol_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| HLE Diamond | 33.8% | Defaultgpt-6-sol | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 6 (Scale rank accounts for CI); ±2.9; entry added 2026-04-08 |
| Humanity's Last Exam | 34.9% | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-low) |
| Humanity's Last Exam | 41.0% | MediumGPT-6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-medium) |
| Humanity's Last Exam | 44.1% | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-high) |
| Humanity's Last Exam | 46.3% | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-xhigh) |
| Humanity's Last Exam | 47.9% | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol) |
| IOI (Vals) | 82.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±9.24 stderr; $2.703/test |
| LMArena Code Arena (WebDev) | 1689 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 7 (CI rank 5-10); 95% CI 1678-1701; 3197 votes |
| LMArena Text - Coding category | 1523 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 28 (CI rank 7-73); 95% CI 1508-1537; 1799 votes |
| LMArena Text - Creative Writing | 1438 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 57 (rank range 24-88); 95% CI 1422.1-1454.1; 1532 votes; style-controlled |
| LMArena Text - Expert | 1498 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 43 (rank range 9-93); 95% CI 1477.9-1518.9; 818 votes; style-controlled |
| LMArena Text - Hard Prompts | 1484 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 55 (rank range 31-80); 95% CI 1475.3-1493.6; 4639 votes; style-controlled |
| LMArena Text - Instruction Following | 1453 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 57 (rank range 32-84); 95% CI 1440.8-1464.6; 2514 votes; style-controlled |
| LMArena Text - Longer Query | 1466 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 61 (rank range 35-92); 95% CI 1455.4-1477.3; 3182 votes; style-controlled |
| LMArena Text - Multi-Turn | 1469 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 61 (rank range 16-101); 95% CI 1450.8-1486.5; 1148 votes; style-controlled |
| LMArena Text - Non-English | 1446 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 60 (rank range 41-86); 95% CI 1436.4-1454.6; 4644 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1448 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 77 (rank range 36-118); 95% CI 1432.1-1464.0; 1435 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1458 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 80 (rank range 20-143); 95% CI 1434.2-1481.3; 649 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1440 | Maxgpt-6-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 69 (rank range 31-98); 95% CI 1425.8-1453.4; 2023 votes; style-controlled |
| OSWorld 2.0 offline set | 43.9% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.0117 |
| OSWorld 2.0 offline set | 54.0% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.3789 |
| OSWorld 2.0 offline set | 58.3% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.7096 |
| OSWorld 2.0 offline set | 60.5% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $2.3002 |
| OSWorld 2.0 offline set | 64.4% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $3.3695 |
| PRBench Legal (Scale) | 38.6% | Defaultgpt-6-sol | Independent testIndependentScale AI (SEAL) ↗ | 25 Sep 2026 | — | Scale rank 28; ±1.58 CI |
| ProgramBench (fully resolved) | 2.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.992 stderr; $11.673/test; strict fully-resolved rate |
| SciCode | 50.2% | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-low) |
| SciCode | 53.8% | MediumGPT-6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-medium) |
| SciCode | 54.9% | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-high) |
| SciCode | 55.1% | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-xhigh) |
| SciCode | 57.6% | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol) |
| SimpleBench | 73.1% | DefaultGPT-6 Sol | Independent testIndependentSimpleBench ↗ | 24 Sep 2026 | — | AVG@5, temp 0.7; rank 15th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 60.7% | Maxgpt-6-sol_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | Epoch-run (no tools); ±1.55 stderr |
| SWE Atlas - Codebase QnA | 57.5% | MaxGPT-6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| Terminal-Bench 2.1 | 83.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.297 stderr; $0.373/test |
| Terminal-Bench 4.0 | 9.1% | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-low) |
| Terminal-Bench 4.0 | 18.7% | MediumGPT-6 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-medium) |
| Terminal-Bench 4.0 | 26.3% | HighGPT-6 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-high) |
| Terminal-Bench 4.0 | 30.3% | Extra highGPT-6 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol-xhigh) |
| Terminal-Bench 4.0 | 44.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±3.535 stderr; $5.793/test |
| Terminal-Bench 4.0 | 43.9% | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-sol) |
| Terminal-Bench 4.0 | 43.4% | MaxGPT-6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench-Science 0.1 | 9.2% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $2.9986 |
| Terminal-Bench-Science 0.1 | 14.5% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $4.4068 |
| Terminal-Bench-Science 0.1 | 14.6% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $4.6325 |
| Terminal-Bench-Science 0.1 | 25.3% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $6.7698 |
| Terminal-Bench-Science 0.1 | 27.6% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $12.1803 |
| Vals Code Migration | 57.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.213 stderr; $15.677/test |
| Vals Index | 57.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.012 stderr; $7.581/test |
| Vals Legal Research Bench | 28.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-6-sol; rank 42/72; ±3.149 stderr; $4.897717/test |
| Vals Public Benefits Bench v1.1 | 56.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-6-sol; rank 35/45; ±1.289 stderr; $2.491613/test |
| Vals Vibe Code Bench | 87.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±2.526 stderr; $26.356/test |
| Vectara Hallucination Leaderboard (HHEM) | 6.5% | Defaultopenai/gpt-6-sol | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 93.5 %; answer rate 100.0 %; avg summary 71.4 words; HHEM-2.3 judge; effort not stated (API default) |
| Vending-Bench 2 | $14,428 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$1,051; rank 2; only top 10 rendered server-side (57 more behind 'Show more') |