Skip to content
Bencher

Models · OpenAI · Out sinceReleased 29 Sep 2026

GPT-6.1 Sol#4 for coding.#4 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-6.1 Sol is made by OpenAI. Among the models we track it ranks #4 for coding, #11 for research and analysis. It's mid-priced to use.

Default reasoning.effort medium; none/minimal not supported. Cached input $0.10/M (95% off). Max input 922K. Prompts >272K input billed 2x input / 1.5x output. reasoning.mode "pro" available (see -pro variant). Knowledge cutoff Apr 30, 2026. OpenAI Codex docs recommend it for complex coding and agentic workflows.

Writing & creativity
—
Research & analysis
73.8 / 100 · #11
Coding
83.7 / 100 · #4
Price
Mid-priced$2 / $10
Price per 1M (blended)Blended / 1M
$4
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
120 (95 independent95 indep.)
Out sinceReleased
29 Sep 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-6.1-sol.

Route it as openai/gpt-6.1-sol at $2 in / $10 out per 1M tokens, 1.1M context. Listed since 29 Sep 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-6.1 SolWhere GPT-6.1 Solthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Artificial Analysis Coding Agent Index

Best at Extra high, worse abovePeaks at Extra high

Terminal-Bench 4.0

Best at MaxBest at Max

DeepSWE

Best at Extra high, worse abovePeaks at Extra high

SWE Atlas - Codebase QnA

Best at Extra high, worse abovePeaks at Extra high

SciCode

Best at High, worse abovePeaks at High

Artificial Analysis Intelligence Index

Best at MaxBest at Max

Terminal-Bench-Science 0.1

Best at MaxBest at Max

AutomationBench

Best at MaxBest at Max

OSWorld 2.0 offline set

Best at MaxBest at Max

AA-LCR

Best at Low, worse abovePeaks at Low

AA-Omniscience Accuracy

Best at MaxBest at Max

AA-Omniscience Hallucination Rate

Best at High, worse abovePeaks at High

Lower is better on this test

AA-Omniscience Index

Best at MaxBest at Max

ARC-AGI-2

Best at MaxBest at Max

ARC-AGI-3

Best at Extra high, worse abovePeaks at Extra high

GDP.pdf

Best at High, worse abovePeaks at High

Humanity's Last Exam

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11564MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR84.0%LowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.3%MediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR82.3%HighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR79.7%Extra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.0%MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy58.9%LowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy60.4%MediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy60.8%HighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy60.8%Extra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy62.1%MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate51.6%LowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate51.6%MediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate49.4%HighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate50.9%Extra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate54.3%MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index37.6LowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index40MediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index41.5HighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index40.9Extra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index41.5MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
ARC-AGI-286.7%MediumGPT-6.1 Sol (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.10/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-291.7%HighGPT-6.1 Sol (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.13/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-291.7%Extra highGPT-6.1 Sol (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.18/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-294.2%MaxGPT-6.1 Sol (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.25/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-382.8%LowGPT-6.1 Sol - Provider Adapter (Low)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $6,622; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-33.9%LowGPT-6.1 Sol (Low)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $7,111; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-391.0%MediumGPT-6.1 Sol - Provider Adapter (Medium)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $5,835; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-310.6%MediumGPT-6.1 Sol (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $7,370; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-395.0%HighGPT-6.1 Sol - Provider Adapter (High)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $5,404; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-326.7%HighGPT-6.1 Sol (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $9,758; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-396.4%Extra highGPT-6.1 Sol - Provider Adapter (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $4,360; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-339.9%Extra highGPT-6.1 Sol (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $9,104; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-396.2%MaxGPT-6.1 Sol - Provider Adapter (Max)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $3,817; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-352.7%MaxGPT-6.1 Sol (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $7,632; data https://arcprize.org/media/data/leaderboard/v3.json
Artificial Analysis Coding Agent Index57.2LowGPT-6.1 Sol (low)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 67.6, SWE-Atlas-QnA 55.1, Terminal-Bench v4 49.0; avg cost $0.50/task; avg wall time 9 min/task
Artificial Analysis Coding Agent Index61.4MediumGPT-6.1 Sol (medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 72.0, SWE-Atlas-QnA 60.8, Terminal-Bench v4 51.5; avg cost $0.70/task; avg wall time 11 min/task
Artificial Analysis Coding Agent Index60.1HighGPT-6.1 Sol (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 70.5, SWE-Atlas-QnA 59.9, Terminal-Bench v4 50.0; avg cost $0.89/task; avg wall time 13 min/task
Artificial Analysis Coding Agent Index62.9Extra highGPT-6.1 Sol (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 73.2, SWE-Atlas-QnA 61.0, Terminal-Bench v4 54.5; avg cost $1.04/task; avg wall time 16 min/task
Artificial Analysis Coding Agent Index60.1MaxGPT-6.1 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 69.6, SWE-Atlas-QnA 57.8, Terminal-Bench v4 53.0; avg cost $1.55/task; avg wall time 24 min/task
Artificial Analysis Intelligence Index42.1LowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-1-sol-low; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.13/task
Artificial Analysis Intelligence Index47.8MediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-1-sol-medium; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.21/task
Artificial Analysis Intelligence Index50.2HighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-1-sol-high; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.32/task
Artificial Analysis Intelligence Index51Extra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-1-sol-xhigh; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.39/task
Artificial Analysis Intelligence Index51.8MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-1-sol; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.72/task
Artificial Analysis output speed55 tok/sLowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 2.7s; list price $2/10 per 1M in/out
Artificial Analysis output speed58 tok/sMediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 5.7s; list price $2/10 per 1M in/out
Artificial Analysis output speed60 tok/sHighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 58.0s; list price $2/10 per 1M in/out
Artificial Analysis output speed63 tok/sExtra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 108.9s; list price $2/10 per 1M in/out
Artificial Analysis output speed64 tok/sMaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 293.4s; list price $2/10 per 1M in/out
AutomationBench24.7%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.157
AutomationBench31.7%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.1917
AutomationBench33.2%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2255
AutomationBench35.5%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2508
AutomationBench36.1%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2989
DeepSWE67.6%LowGPT-6.1 Sol (low)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE64.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.1714
DeepSWE72.0%MediumGPT-6.1 Sol (medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE73.0%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.4196
DeepSWE70.5%HighGPT-6.1 Sol (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE75.2%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.6461
DeepSWE73.2%Extra highGPT-6.1 Sol (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE71.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.7886
DeepSWE69.6%MaxGPT-6.1 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE71.9%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.5711
FrontierCode50.2%Mediumgpt-6.1-sol_mediumIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026codexread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5
FrontierMath (Tiers 1-3)93.7%Maxgpt-6.1-sol_maxIndependent testIndependentEpoch AI ↗29 Sep 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 4100.0%Maxgpt-6.1-sol_maxIndependent testIndependentEpoch AI ↗29 Sep 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 0.0pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
GDP.pdf27.0%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.3341
GDP.pdf30.0%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.3375
GDP.pdf32.0%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.3494
GDP.pdf31.8%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.3681
GDP.pdf31.0%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.4199
GDPval-AA v2.11575MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1554.9-1595.28
GPQA Diamond95.4%Maxgpt-6.1-sol_maxIndependent testIndependentEpoch AI ↗29 Sep 2026—Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
Humanity's Last Exam47.4%LowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-low)
Humanity's Last Exam49.9%MediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-medium)
Humanity's Last Exam51.4%HighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-high)
Humanity's Last Exam52.6%Extra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-xhigh)
Humanity's Last Exam52.9%MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol)
IOI (Vals)96.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±3.111 stderr; $1.262/test
LMArena Code Arena (WebDev)1759Maxgpt-6.1-sol-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 3 (CI rank 3-4); 95% CI 1740-1778; 1264 votes
OSWorld 2.0 offline set59.0%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.4248
OSWorld 2.0 offline set66.8%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.7675
OSWorld 2.0 offline set69.6%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.9613
OSWorld 2.0 offline set69.4%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.0466
OSWorld 2.0 offline set71.4%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.2689
SciCode53.2%LowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-low)
SciCode53.2%MediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-medium)
SciCode55.8%HighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-high)
SciCode55.7%Extra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-xhigh)
SciCode54.2%MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol)
SimpleQA Verified73.9%Maxgpt-6.1-sol_maxIndependent testIndependentEpoch AI ↗29 Sep 2026—Epoch-run (no tools); ±1.39 stderr
SWE Atlas - Codebase QnA55.1%LowGPT-6.1 Sol (low)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA60.8%MediumGPT-6.1 Sol (medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA59.9%HighGPT-6.1 Sol (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA61.0%Extra highGPT-6.1 Sol (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA57.8%MaxGPT-6.1 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
Terminal-Bench 4.049.0%LowGPT-6.1 Sol (low)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.030.8%LowGPT-6.1 Sol (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-low)
Terminal-Bench 4.051.5%MediumGPT-6.1 Sol (medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.048.0%MediumGPT-6.1 Sol (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-medium)
Terminal-Bench 4.051.5%HighGPT-6.1 Sol (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-high)
Terminal-Bench 4.050.0%HighGPT-6.1 Sol (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.054.5%Extra highGPT-6.1 Sol (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.054.0%Extra highGPT-6.1 Sol (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol-xhigh)
Terminal-Bench 4.056.1%MaxGPT-6.1 Sol (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-1-sol)
Terminal-Bench 4.055.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.821 stderr; $1.725/test
Terminal-Bench 4.053.0%MaxGPT-6.1 Sol (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench-Science 0.143.7%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $1.7918
Terminal-Bench-Science 0.147.6%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $2.3386
Terminal-Bench-Science 0.151.1%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $2.7594
Terminal-Bench-Science 0.153.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $2.8922
Terminal-Bench-Science 0.157.0%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $5.4652
Vals Code Migration65.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.362 stderr; $6.506/test
Vals Index61.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.011 stderr; $3.237/test
Vals Legal Research Bench38.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-6.1-sol; rank 30/72; ±3.381 stderr; $3.384561/test
Vals Public Benefits Bench v1.159.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-6.1-sol; rank 31/45; ±1.278 stderr; $0.690795/test
Vals SRE Bench50.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±3.095 stderr; $2.685/test
Vals Vibe Code Bench88.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±1.893 stderr; $6.235/test