Skip to content
Bencher

Models · OpenAI · Out sinceReleased 22 Sep 2026

GPT-6 Sol#10 for coding.#10 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-6 Sol is made by OpenAI. Among the models we track it ranks #10 for coding, #28 for writing, #31 for research and analysis. It's mid-priced to use.

Superseded by GPT-6.1 Sol a week later. Default effort medium. Cached input $0.20/M. Knowledge cutoff Apr 20, 2026. System card link on launch post points to the GPT-6 Astra system card.

Writing & creativity
36.5 / 100 · #28
Research & analysis
49.7 / 100 · #31
Coding
64.0 / 100 · #10
Price
Mid-priced$2 / $10
Price per 1M (blended)Blended / 1M
$4
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
138 (103 independent103 indep.)
Out sinceReleased
22 Sep 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-6-sol.

Route it as openai/gpt-6-sol at $2 in / $10 out per 1M tokens, 1.1M context. Listed since 22 Sep 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-6 SolWhere GPT-6 Solthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MaxBest at Max

DeepSWE

Best at MaxBest at Max

FrontierCode

Best at MaxBest at Max

SciCode

Best at MaxBest at Max

Agents' Last Exam

Best at MaxBest at Max

Artificial Analysis Intelligence Index

Best at MaxBest at Max

Terminal-Bench-Science 0.1

Best at MaxBest at Max

AutomationBench

Best at Extra high, worse abovePeaks at Extra high

OSWorld 2.0 offline set

Best at MaxBest at Max

AA-Briefcase v1.1

Best at MaxBest at Max

AA-LCR

Best at HighBest at High

AA-Omniscience Accuracy

Best at MaxBest at Max

AA-Omniscience Hallucination Rate

Best at Low, worse abovePeaks at Low

Lower is better on this test

AA-Omniscience Index

Best at MaxBest at Max

ARC-AGI-3

Best at MaxBest at Max

GDP.pdf

Best at High, worse abovePeaks at High

Humanity's Last Exam

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.1905Lowreasoning effort=lowIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $0.12
AA-Briefcase v1.11142Mediumreasoning effort=medIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $0.34
AA-Briefcase v1.11289Highreasoning effort=highIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $0.63
AA-Briefcase v1.11364Extra highreasoning effort=xhighIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $1.19
AA-Briefcase v1.11483Maxreasoning effort=maxIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); cost/task $2.67
AA-Briefcase v1.11479MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR79.3%LowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR82.3%MediumGPT-6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.7%HighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR81.3%Extra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.7%MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy51.3%LowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy53.5%MediumGPT-6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy53.7%HighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy53.9%Extra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy54.5%MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate50.7%LowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate56.8%MediumGPT-6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate58.1%HighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate58.9%Extra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate60.1%MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index26.5LowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index27MediumGPT-6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index26.8HighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index26.7Extra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index27.1MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam48.7%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.8641
Agents' Last Exam53.1%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.2664
Agents' Last Exam52.6%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.5302
Agents' Last Exam55.4%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.668
Agents' Last Exam56.4%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $2.9313
ARC-AGI-289.6%MaxGPT-6 Sol (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.44/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-31.2%No reasoningGPT-6 Sol - Provider Adapter (None)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $4,234; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.3%No reasoningGPT-6 Sol (None)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $4,656; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-31.8%LowGPT-6 Sol - Provider Adapter (Low)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $4,745; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-34.9%MediumGPT-6 Sol - Provider Adapter (Medium)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $5,889; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.4%MediumGPT-6 Sol (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $4,864; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-35.9%HighGPT-6 Sol - Provider Adapter (High)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $5,908; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.9%HighGPT-6 Sol (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $5,072; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-39.5%Extra highGPT-6 Sol - Provider Adapter (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $6,907; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-31.8%Extra highGPT-6 Sol (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $5,106; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-323.0%MaxGPT-6 Sol - Provider Adapter (Max)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $8,722; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-34.6%MaxGPT-6 Sol (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $5,554; data https://arcprize.org/media/data/leaderboard/v3.json
Artificial Analysis Coding Agent Index56.7MaxGPT-6 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 69.0, SWE-Atlas-QnA 57.5, Terminal-Bench v4 43.4; avg cost $2.99/task; avg wall time 22 min/task
Artificial Analysis Intelligence Index34.2LowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-sol-low; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.13/task
Artificial Analysis Intelligence Index39.8MediumGPT-6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-sol-medium; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.25/task
Artificial Analysis Intelligence Index42.4HighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-sol-high; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.37/task
Artificial Analysis Intelligence Index44.2Extra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-sol-xhigh; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.52/task
Artificial Analysis Intelligence Index47.6MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-sol; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $1.04/task
Artificial Analysis output speed56 tok/sLowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 1.9s; list price $2/10 per 1M in/out
Artificial Analysis output speed63 tok/sHighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 19.2s; list price $2/10 per 1M in/out
Artificial Analysis output speed68 tok/sExtra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 57.8s; list price $2/10 per 1M in/out
Artificial Analysis output speed74 tok/sMaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 177.8s; list price $2/10 per 1M in/out
AutomationBench21.2%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.1861
AutomationBench26.9%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2098
AutomationBench31.2%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2368
AutomationBench33.2%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2746
AutomationBench32.0%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.3406
Chartography (no tools)53.6%Defaultas listed by Anthropic (effort not stated)Independent testIndependentSurge AI (via Anthropic launch post) ↗28 Sep 2026—no tools
DeepSWE37.2%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.1623
DeepSWE56.6%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.3798
DeepSWE65.3%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.6404
DeepSWE66.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.0033
DeepSWE69.0%MaxGPT-6 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE68.8%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $2.7439
Design Arena (all categories)1292Defaultgpt-6-solIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±7.8 SE; 2044 battles; win rate 52%
Design Arena (fullstack)1237Defaultgpt-6-solIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±10.8 SE; 1154 battles; win rate 52.2%
EQ-Bench Creative Writing v3 (Elo)2125Default*gpt-6-solIndependent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 4; rubric score 16.51/20; slop 8.97; avg length 6182 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*)
FrontierCode37.3%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.453
FrontierCode45.9%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.7966
FrontierCode47.7%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $1.0793
FrontierCode48.5%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $1.3748
FrontierCode49.3%Maxgpt-6-sol_maxIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026codexread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5
FrontierCode49.3%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.137
FrontierMath (Tiers 1-3)89.8%Maxgpt-6-sol_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 490.0%Maxgpt-6-sol_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 5.2pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
GDP.pdf21.8%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.33
GDP.pdf25.4%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.34
GDP.pdf28.0%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.35
GDP.pdf23.8%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.37
GDP.pdf24.8%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.43
GDPval-AA v2.11505MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1480.19-1530.25
GDPval-AA v2.11487Defaultas listed by Anthropic (effort not stated)Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗28 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); may predate OpenAI image-understanding bug fix for GPT-6 Sol
GPQA Diamond94.3%Maxgpt-6-sol_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
HLE Diamond33.8%Defaultgpt-6-solIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 6 (Scale rank accounts for CI); ±2.9; entry added 2026-04-08
Humanity's Last Exam34.9%LowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-low)
Humanity's Last Exam41.0%MediumGPT-6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-medium)
Humanity's Last Exam44.1%HighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-high)
Humanity's Last Exam46.3%Extra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-xhigh)
Humanity's Last Exam47.9%MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol)
IOI (Vals)82.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±9.24 stderr; $2.703/test
LMArena Code Arena (WebDev)1689Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 7 (CI rank 5-10); 95% CI 1678-1701; 3197 votes
LMArena Text - Coding category1523Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 28 (CI rank 7-73); 95% CI 1508-1537; 1799 votes
LMArena Text - Creative Writing1438Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 57 (rank range 24-88); 95% CI 1422.1-1454.1; 1532 votes; style-controlled
LMArena Text - Expert1498Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 43 (rank range 9-93); 95% CI 1477.9-1518.9; 818 votes; style-controlled
LMArena Text - Hard Prompts1484Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 55 (rank range 31-80); 95% CI 1475.3-1493.6; 4639 votes; style-controlled
LMArena Text - Instruction Following1453Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 57 (rank range 32-84); 95% CI 1440.8-1464.6; 2514 votes; style-controlled
LMArena Text - Longer Query1466Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 61 (rank range 35-92); 95% CI 1455.4-1477.3; 3182 votes; style-controlled
LMArena Text - Multi-Turn1469Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 61 (rank range 16-101); 95% CI 1450.8-1486.5; 1148 votes; style-controlled
LMArena Text - Non-English1446Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 60 (rank range 41-86); 95% CI 1436.4-1454.6; 4644 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1448Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 77 (rank range 36-118); 95% CI 1432.1-1464.0; 1435 votes; style-controlled
LMArena Text - Occupational: Legal & Government1458Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 80 (rank range 20-143); 95% CI 1434.2-1481.3; 649 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1440Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 69 (rank range 31-98); 95% CI 1425.8-1453.4; 2023 votes; style-controlled
OSWorld 2.0 offline set43.9%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.0117
OSWorld 2.0 offline set54.0%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.3789
OSWorld 2.0 offline set58.3%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.7096
OSWorld 2.0 offline set60.5%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $2.3002
OSWorld 2.0 offline set64.4%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $3.3695
PRBench Legal (Scale)38.6%Defaultgpt-6-solIndependent testIndependentScale AI (SEAL) ↗25 Sep 2026—Scale rank 28; ±1.58 CI
ProgramBench (fully resolved)2.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.992 stderr; $11.673/test; strict fully-resolved rate
SciCode50.2%LowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-low)
SciCode53.8%MediumGPT-6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-medium)
SciCode54.9%HighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-high)
SciCode55.1%Extra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-xhigh)
SciCode57.6%MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol)
SimpleBench73.1%DefaultGPT-6 SolIndependent testIndependentSimpleBench ↗24 Sep 2026—AVG@5, temp 0.7; rank 15th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified60.7%Maxgpt-6-sol_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—Epoch-run (no tools); ±1.55 stderr
SWE Atlas - Codebase QnA57.5%MaxGPT-6 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
Terminal-Bench 2.183.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.297 stderr; $0.373/test
Terminal-Bench 4.09.1%LowGPT-6 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-low)
Terminal-Bench 4.018.7%MediumGPT-6 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-medium)
Terminal-Bench 4.026.3%HighGPT-6 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-high)
Terminal-Bench 4.030.3%Extra highGPT-6 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol-xhigh)
Terminal-Bench 4.044.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±3.535 stderr; $5.793/test
Terminal-Bench 4.043.9%MaxGPT-6 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-sol)
Terminal-Bench 4.043.4%MaxGPT-6 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench-Science 0.19.2%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $2.9986
Terminal-Bench-Science 0.114.5%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $4.4068
Terminal-Bench-Science 0.114.6%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $4.6325
Terminal-Bench-Science 0.125.3%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $6.7698
Terminal-Bench-Science 0.127.6%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $12.1803
Vals Code Migration57.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.213 stderr; $15.677/test
Vals Index57.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.012 stderr; $7.581/test
Vals Legal Research Bench28.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-6-sol; rank 42/72; ±3.149 stderr; $4.897717/test
Vals Public Benefits Bench v1.156.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-6-sol; rank 35/45; ±1.289 stderr; $2.491613/test
Vals Vibe Code Bench87.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±2.526 stderr; $26.356/test
Vectara Hallucination Leaderboard (HHEM)6.5%Defaultopenai/gpt-6-solIndependent testIndependentVectara ↗22 Sep 2026—factual consistency 93.5 %; answer rate 100.0 %; avg summary 71.4 words; HHEM-2.3 judge; effort not stated (API default)
Vending-Bench 2$14,428DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$1,051; rank 2; only top 10 rendered server-side (57 more behind 'Show more')