Models · OpenAI · Out sinceReleased 29 Sep 2026
GPT-6.1 Sol#4 for coding.#4 for coding, best at max effort.
GPT-6.1 Sol is made by OpenAI. Among the models we track it ranks #4 for coding, #11 for research and analysis. It's mid-priced to use.
Default reasoning.effort medium; none/minimal not supported. Cached input $0.10/M (95% off). Max input 922K. Prompts >272K input billed 2x input / 1.5x output. reasoning.mode "pro" available (see -pro variant). Knowledge cutoff Apr 30, 2026. OpenAI Codex docs recommend it for complex coding and agentic workflows.
- —
- 73.8 / 100 · #11
- 83.7 / 100 · #4
- Mid-priced$2 / $10
- $4
- 1.1M
- 128K
- 120 (95 independent95 indep.)
- 29 Sep 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-6.1-sol.
Route it as openai/gpt-6.1-sol at $2 in / $10 out per 1M tokens, 1.1M context. Listed since 29 Sep 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-6.1 SolWhere GPT-6.1 Solthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Artificial Analysis Coding Agent Index
Best at Extra high, worse abovePeaks at Extra high
Terminal-Bench 4.0
Best at MaxBest at Max
DeepSWE
Best at Extra high, worse abovePeaks at Extra high
SWE Atlas - Codebase QnA
Best at Extra high, worse abovePeaks at Extra high
SciCode
Best at High, worse abovePeaks at High
Artificial Analysis Intelligence Index
Best at MaxBest at Max
Terminal-Bench-Science 0.1
Best at MaxBest at Max
AutomationBench
Best at MaxBest at Max
OSWorld 2.0 offline set
Best at MaxBest at Max
AA-LCR
Best at Low, worse abovePeaks at Low
AA-Omniscience Accuracy
Best at MaxBest at Max
AA-Omniscience Hallucination Rate
Best at High, worse abovePeaks at High
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
ARC-AGI-2
Best at MaxBest at Max
ARC-AGI-3
Best at Extra high, worse abovePeaks at Extra high
GDP.pdf
Best at High, worse abovePeaks at High
Humanity's Last Exam
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1564 | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 84.0% | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.3% | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 82.3% | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 79.7% | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.0% | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 58.9% | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 60.4% | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 60.8% | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 60.8% | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 62.1% | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 51.6% | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 51.6% | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 49.4% | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 50.9% | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 54.3% | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 37.6 | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 40 | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 41.5 | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 40.9 | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 41.5 | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| ARC-AGI-2 | 86.7% | MediumGPT-6.1 Sol (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.10/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 91.7% | HighGPT-6.1 Sol (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.13/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 91.7% | Extra highGPT-6.1 Sol (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.18/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 94.2% | MaxGPT-6.1 Sol (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.25/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-3 | 82.8% | LowGPT-6.1 Sol - Provider Adapter (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $6,622; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 3.9% | LowGPT-6.1 Sol (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $7,111; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 91.0% | MediumGPT-6.1 Sol - Provider Adapter (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $5,835; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 10.6% | MediumGPT-6.1 Sol (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $7,370; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 95.0% | HighGPT-6.1 Sol - Provider Adapter (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $5,404; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 26.7% | HighGPT-6.1 Sol (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $9,758; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 96.4% | Extra highGPT-6.1 Sol - Provider Adapter (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $4,360; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 39.9% | Extra highGPT-6.1 Sol (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $9,104; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 96.2% | MaxGPT-6.1 Sol - Provider Adapter (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $3,817; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 52.7% | MaxGPT-6.1 Sol (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $7,632; data https://arcprize.org/media/data/leaderboard/v3.json |
| Artificial Analysis Coding Agent Index | 57.2 | LowGPT-6.1 Sol (low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 67.6, SWE-Atlas-QnA 55.1, Terminal-Bench v4 49.0; avg cost $0.50/task; avg wall time 9 min/task |
| Artificial Analysis Coding Agent Index | 61.4 | MediumGPT-6.1 Sol (medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 72.0, SWE-Atlas-QnA 60.8, Terminal-Bench v4 51.5; avg cost $0.70/task; avg wall time 11 min/task |
| Artificial Analysis Coding Agent Index | 60.1 | HighGPT-6.1 Sol (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 70.5, SWE-Atlas-QnA 59.9, Terminal-Bench v4 50.0; avg cost $0.89/task; avg wall time 13 min/task |
| Artificial Analysis Coding Agent Index | 62.9 | Extra highGPT-6.1 Sol (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 73.2, SWE-Atlas-QnA 61.0, Terminal-Bench v4 54.5; avg cost $1.04/task; avg wall time 16 min/task |
| Artificial Analysis Coding Agent Index | 60.1 | MaxGPT-6.1 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 69.6, SWE-Atlas-QnA 57.8, Terminal-Bench v4 53.0; avg cost $1.55/task; avg wall time 24 min/task |
| Artificial Analysis Intelligence Index | 42.1 | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-1-sol-low; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.13/task |
| Artificial Analysis Intelligence Index | 47.8 | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-1-sol-medium; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.21/task |
| Artificial Analysis Intelligence Index | 50.2 | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-1-sol-high; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.32/task |
| Artificial Analysis Intelligence Index | 51 | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-1-sol-xhigh; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.39/task |
| Artificial Analysis Intelligence Index | 51.8 | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-1-sol; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.72/task |
| Artificial Analysis output speed | 55 tok/s | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 2.7s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 58 tok/s | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 5.7s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 60 tok/s | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 58.0s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 63 tok/s | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 108.9s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 64 tok/s | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 293.4s; list price $2/10 per 1M in/out |
| AutomationBench | 24.7% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.157 |
| AutomationBench | 31.7% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.1917 |
| AutomationBench | 33.2% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2255 |
| AutomationBench | 35.5% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2508 |
| AutomationBench | 36.1% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.2989 |
| DeepSWE | 67.6% | LowGPT-6.1 Sol (low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 64.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.1714 |
| DeepSWE | 72.0% | MediumGPT-6.1 Sol (medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 73.0% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.4196 |
| DeepSWE | 70.5% | HighGPT-6.1 Sol (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 75.2% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.6461 |
| DeepSWE | 73.2% | Extra highGPT-6.1 Sol (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 71.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.7886 |
| DeepSWE | 69.6% | MaxGPT-6.1 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 71.9% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.5711 |
| FrontierCode | 50.2% | Mediumgpt-6.1-sol_medium | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | codex | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5 |
| FrontierMath (Tiers 1-3) | 93.7% | Maxgpt-6.1-sol_max | Independent testIndependentEpoch AI ↗ | 29 Sep 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 100.0% | Maxgpt-6.1-sol_max | Independent testIndependentEpoch AI ↗ | 29 Sep 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 0.0pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| GDP.pdf | 27.0% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.3341 |
| GDP.pdf | 30.0% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.3375 |
| GDP.pdf | 32.0% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.3494 |
| GDP.pdf | 31.8% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.3681 |
| GDP.pdf | 31.0% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.4199 |
| GDPval-AA v2.1 | 1575 | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1554.9-1595.28 |
| GPQA Diamond | 95.4% | Maxgpt-6.1-sol_max | Independent testIndependentEpoch AI ↗ | 29 Sep 2026 | — | Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| Humanity's Last Exam | 47.4% | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-low) |
| Humanity's Last Exam | 49.9% | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-medium) |
| Humanity's Last Exam | 51.4% | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-high) |
| Humanity's Last Exam | 52.6% | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-xhigh) |
| Humanity's Last Exam | 52.9% | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol) |
| IOI (Vals) | 96.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±3.111 stderr; $1.262/test |
| LMArena Code Arena (WebDev) | 1759 | Maxgpt-6.1-sol-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 3 (CI rank 3-4); 95% CI 1740-1778; 1264 votes |
| OSWorld 2.0 offline set | 59.0% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.4248 |
| OSWorld 2.0 offline set | 66.8% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.7675 |
| OSWorld 2.0 offline set | 69.6% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.9613 |
| OSWorld 2.0 offline set | 69.4% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.0466 |
| OSWorld 2.0 offline set | 71.4% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $1.2689 |
| SciCode | 53.2% | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-low) |
| SciCode | 53.2% | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-medium) |
| SciCode | 55.8% | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-high) |
| SciCode | 55.7% | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-xhigh) |
| SciCode | 54.2% | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol) |
| SimpleQA Verified | 73.9% | Maxgpt-6.1-sol_max | Independent testIndependentEpoch AI ↗ | 29 Sep 2026 | — | Epoch-run (no tools); ±1.39 stderr |
| SWE Atlas - Codebase QnA | 55.1% | LowGPT-6.1 Sol (low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 60.8% | MediumGPT-6.1 Sol (medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 59.9% | HighGPT-6.1 Sol (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 61.0% | Extra highGPT-6.1 Sol (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 57.8% | MaxGPT-6.1 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 49.0% | LowGPT-6.1 Sol (low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 30.8% | LowGPT-6.1 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-low) |
| Terminal-Bench 4.0 | 51.5% | MediumGPT-6.1 Sol (medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 48.0% | MediumGPT-6.1 Sol (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-medium) |
| Terminal-Bench 4.0 | 51.5% | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-high) |
| Terminal-Bench 4.0 | 50.0% | HighGPT-6.1 Sol (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 54.5% | Extra highGPT-6.1 Sol (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 54.0% | Extra highGPT-6.1 Sol (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol-xhigh) |
| Terminal-Bench 4.0 | 56.1% | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-1-sol) |
| Terminal-Bench 4.0 | 55.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.821 stderr; $1.725/test |
| Terminal-Bench 4.0 | 53.0% | MaxGPT-6.1 Sol (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench-Science 0.1 | 43.7% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $1.7918 |
| Terminal-Bench-Science 0.1 | 47.6% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $2.3386 |
| Terminal-Bench-Science 0.1 | 51.1% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $2.7594 |
| Terminal-Bench-Science 0.1 | 53.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $2.8922 |
| Terminal-Bench-Science 0.1 | 57.0% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $5.4652 |
| Vals Code Migration | 65.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.362 stderr; $6.506/test |
| Vals Index | 61.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.011 stderr; $3.237/test |
| Vals Legal Research Bench | 38.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-6.1-sol; rank 30/72; ±3.381 stderr; $3.384561/test |
| Vals Public Benefits Bench v1.1 | 59.3% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-6.1-sol; rank 31/45; ±1.278 stderr; $0.690795/test |
| Vals SRE Bench | 50.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±3.095 stderr; $2.685/test |
| Vals Vibe Code Bench | 88.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±1.893 stderr; $6.235/test |