Skip to content
Bencher

Models · OpenAI · Out sinceReleased 3 Sep 2026

GPT-6 Astra#3 for coding.#3 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)ProprietaryOur pick: Best for the hardest thinking problemsPick: Best for hard reasoning
Where you can use it:In apps:ChatGPT (Business plan) · Pro (GPT-6 Pro)

GPT-6 Astra is made by OpenAI. Among the models we track it ranks #3 for coding, #12 for writing, #13 for research and analysis. It's expensive to use. It solves the toughest logic and science puzzles better than any other model.

Flagship. Announced Sep 3, 2026 (limited orgs), API in following days (OpenRouter listing Sep 4). none effort not supported. Cached input $1/M, cache writes $12.5/M. >272K input: 2x input, 1.5x output. Fast mode 2x price; Ultrafast mode available. Meets OpenAI Preparedness "Critical" cyber threshold; misalignment monitoring in production. Knowledge cutoff Apr 30, 2026.

Writing & creativity
57.7 / 100 · #12
Research & analysis
72.3 / 100 · #13
Coding
85.4 / 100 · #3
Price
Expensive$10 / $50
Price per 1M (blended)Blended / 1M
$20
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
220 (152 independent152 indep.)
Out sinceReleased
3 Sep 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-6-astra.

Route it as openai/gpt-6-astra at $10 in / $50 out per 1M tokens, 1.1M context. Listed since 4 Sep 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-6 AstraWhere GPT-6 Astrathink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Artificial Analysis Coding Agent Index

Best at High, worse abovePeaks at High

Terminal-Bench 4.0

Best at Extra high, worse abovePeaks at Extra high

DeepSWE

Best at Extra high, worse abovePeaks at Extra high

FrontierCode

Best at MaxBest at Max

SWE Atlas - Codebase QnA

Best at MaxBest at Max

Terminal-Bench 2.1

Best at High, worse abovePeaks at High

SciCode

Best at MaxBest at Max

Agents' Last Exam

Best at MaxBest at Max

Artificial Analysis Intelligence Index

Best at MaxBest at Max

Terminal-Bench-Science 0.1

Best at MaxBest at Max

AutomationBench

Best at MaxBest at Max

OSWorld 2.0 offline set

Best at MaxBest at Max

AA-LCR

Best at MaxBest at Max

AA-Omniscience Accuracy

Best at MaxBest at Max

AA-Omniscience Hallucination Rate

Best at High, worse abovePeaks at High

Lower is better on this test

AA-Omniscience Index

Best at High, worse abovePeaks at High

ARC-AGI-2

Best at MaxBest at Max

ARC-AGI-3

Best at High, worse abovePeaks at High

FrontierMath Tier 4

Best at MediumBest at Medium

GDP.pdf

Best at Extra high, worse abovePeaks at Extra high

GDPval-AA v2.1

Best at MaxBest at Max

GPQA Diamond

Best at Extra high, worse abovePeaks at Extra high

Humanity's Last Exam

Best at MaxBest at Max

ScreenSpot-Pro

Best at Extra high, worse abovePeaks at Extra high

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent51.3%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase v1.11569MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR80.0%LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR79.7%MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR80.0%HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR80.0%Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR80.7%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy59.5%LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy60.6%MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy61.1%HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy61.9%Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy62.6%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate46.9%LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate46.5%MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate44.8%HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate48.3%Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate51.3%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index40.6LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index42.2MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index43.7HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index43.4Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index43.4MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam53.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $3.0726
Agents' Last Exam57.6%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $4.1046
Agents' Last Exam57.8%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $4.6396
Agents' Last Exam58.3%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $5.397
Agents' Last Exam59.3%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $6.2337
ARC-AGI-198.5%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
ARC-AGI-285.4%LowGPT-6 Astra (Low)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.42/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-292.1%MediumGPT-6 Astra (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.48/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-292.1%HighGPT-6 Astra (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.67/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-293.3%Extra highGPT-6 Astra (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.83/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-295.0%MaxGPT-6 Astra (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.12/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-295.0%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
ARC-AGI-396.7%No reasoningGPT-6 Astra - Provider Adapter (None)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $23,457; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-335.2%No reasoningGPT-6 Astra (None)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $49,791; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-398.0%LowGPT-6 Astra - Provider Adapter (Low)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $21,298; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-317.5%LowGPT-6 Astra (Low)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $38,166; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-398.4%MediumGPT-6 Astra - Provider Adapter (Medium)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $19,285; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-338.6%MediumGPT-6 Astra (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $48,090; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-399.9%HighGPT-6 Astra - Provider Adapter (High)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $18,817; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-354.8%HighGPT-6 Astra (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $40,705; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-398.4%Extra highGPT-6 Astra - Provider Adapter (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $18,147; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-359.3%Extra highGPT-6 Astra (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $37,317; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-398.6%MaxGPT-6 Astra - Provider Adapter (Max)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $17,332; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-362.7%MaxGPT-6 Astra (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $26,098; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-399.9%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—run with Responses API harness (two settings changed)
Artificial Analysis Coding Agent Index62.6Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $1.50329
Artificial Analysis Coding Agent Index65.3Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $2.32706
Artificial Analysis Coding Agent Index65.5Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.06706
Artificial Analysis Coding Agent Index58.9Extra highGPT-6 Astra XHigh + SWE-2 MediumIndependent testIndependentArtificial Analysis ↗1 Oct 2026Devin Fusion CLIagent Devin Fusion CLI; components: DeepSWE v1.1 67.3, SWE-Atlas-QnA 59.4, Terminal-Bench v4 50.0; avg cost $4.54/task; avg wall time 25 min/task; Devin Fusion = Fable/GPT planner + Cognition SWE-2 (medium) sidekick
Artificial Analysis Coding Agent Index67Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.46496
Artificial Analysis Coding Agent Index61.6MaxGPT-6 Astra (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 67.6, SWE-Atlas-QnA 61.8, Terminal-Bench v4 55.6; avg cost $7.47/task; avg wall time 29 min/task
Artificial Analysis Coding Agent Index67Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $4.9813
Artificial Analysis Intelligence Index45.8LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-astra-low; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $0.82/task
Artificial Analysis Intelligence Index49.6MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-astra-medium; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $1.54/task
Artificial Analysis Intelligence Index50.9HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-astra-high; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $1.73/task
Artificial Analysis Intelligence Index52.4Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-astra-xhigh; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $2.31/task
Artificial Analysis Intelligence Index52.7MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-astra; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $3.26/task
Artificial Analysis Intelligence Index61.2Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—Artificial Analysis Intelligence Index v4.1.1
Artificial Analysis output speed43 tok/sLowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 2.9s; list price $10/50 per 1M in/out
Artificial Analysis output speed44 tok/sMediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 6.2s; list price $10/50 per 1M in/out
Artificial Analysis output speed44 tok/sHighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 65.6s; list price $10/50 per 1M in/out
Artificial Analysis output speed48 tok/sExtra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 205.7s; list price $10/50 per 1M in/out
Artificial Analysis output speed51 tok/sMaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 326.7s; list price $10/50 per 1M in/out
AutomationBench30.3%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.08
AutomationBench34.1%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.27
AutomationBench37.1%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.44
AutomationBench39.0%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.5
AutomationBench41.4%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.73
BenchCAD95.9%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—with python tool
BrowseComp91.5%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
DeepSWE67.0%Lowgpt-6-astra_lowIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 79.6%; ±1.3; 4 runs; $2.19/task
DeepSWE67.0%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.5952
DeepSWE72.8%Mediumgpt-6-astra_mediumIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 82.3%; ±2.6; 4 runs; $4.38/task
DeepSWE72.8%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.0755
DeepSWE73.2%Highgpt-6-astra_highIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 82.3%; ±3.4; 4 runs; $5.72/task
DeepSWE73.2%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.9237
DeepSWE74.1%Extra highgpt-6-astra_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 80.5%; ±2.9; 4 runs; $6.52/task
DeepSWE67.3%Extra highGPT-6 Astra XHigh + SWE-2 MediumIndependent testIndependentArtificial Analysis ↗1 Oct 2026Devin Fusion CLIAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE74.1%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $4.4291
DeepSWE73.2%Maxgpt-6-astra_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 79.6%; ±0.8; 4 runs; $12.37/task
DeepSWE67.6%MaxGPT-6 Astra (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE73.2%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $7.4978
Design Arena (all categories)1361Maxgpt-6-astra-maxIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±9.2 SE; 1521 battles; win rate 60.6%
Design Arena (all categories)1388Defaultgpt-6-astraIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±4.6 SE; 6610 battles; win rate 65.2%
Design Arena (fullstack)1355Defaultgpt-6-astraIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±15.5 SE; 553 battles; win rate 61.1%
Epoch Capabilities Index166.5DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 163.1-171.1; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)2173Default*gpt-6-astraIndependent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 1; rubric score 16.80/20; slop 8.41; avg length 6191 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*)
EuroEval Swedish (generative)1.2Defaultgpt-6-astra (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 1; ±0.04; lower is better; task scores (first metric): SweDN summarisation 39.48 ± 0.17, Skolprov 91.00 ± 1.67, Swedish facts 88.39 ± 1.97, ScaLA-sv 83.01 ± 0.94
FrontierCode45.3%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $1.7
FrontierCode48.8%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.43
FrontierCode50.9%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $3.01
FrontierCode50.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $3.28
FrontierCode53.3%Maxgpt-6-astra_maxIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026codexread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5
FrontierCode53.3%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $4.59
FrontierCode v1.1 (Extended)64.5%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—run with a Codex-like developer message discouraging excess tests/unrelated cleanup
FrontierMath (Tiers 1-3)93.7%Maxgpt-6-astra_maxIndependent testIndependentEpoch AI ↗30 Aug 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 482.9%No reasoninggpt-6-astra_noneIndependent testIndependentEpoch AI ↗30 Aug 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 5.9pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 487.8%Lowgpt-6-astra_lowIndependent testIndependentEpoch AI ↗30 Aug 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 5.2pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 497.6%Mediumgpt-6-astra_mediumIndependent testIndependentEpoch AI ↗30 Aug 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 497.6%Highgpt-6-astra_highIndependent testIndependentEpoch AI ↗30 Aug 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 497.6%Extra highgpt-6-astra_xhighIndependent testIndependentEpoch AI ↗30 Aug 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 497.6%Maxgpt-6-astra_maxIndependent testIndependentEpoch AI ↗30 Aug 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 497.6%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—v2
FrontierSWE65.5%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±8.9; $1029.65/trial; 12.1h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort
FrontierSWE v265.5%Maxreasoning effort=maxIndependent testIndependentProximal (via Anthropic system card) ↗22 Sep 2026Proximal harnessGPT-6 Astra ranks first on FrontierSWE v2
GDP.pdf30.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.69632
GDP.pdf30.4%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.71981
GDP.pdf31.0%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.79482
GDP.pdf32.2%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.91272
GDP.pdf31.0%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $2.07582
GDPval-AA v2.11366Lowreasoning effort=lowIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $0.85
GDPval-AA v2.11468Mediumreasoning effort=medIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $1.82
GDPval-AA v2.11485Highreasoning effort=highIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $2.43
GDPval-AA v2.11516Extra highreasoning effort=xhighIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $3.04
GDPval-AA v2.11542Maxreasoning effort=maxIndependent testIndependentArtificial Analysis (via Anthropic launch post) ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $4.53
GDPval-AA v2.11542MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1516.58-1567.2
GPQA Diamond93.1%LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-low)
GPQA Diamond91.8%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.05299
GPQA Diamond93.9%MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-medium)
GPQA Diamond94.9%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.05965
GPQA Diamond94.9%HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-high)
GPQA Diamond93.7%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.07805
GPQA Diamond96.3%Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-xhigh)
GPQA Diamond94.4%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.09368
GPQA Diamond96.1%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra)
GPQA Diamond95.8%Maxgpt-6-astra_maxIndependent testIndependentEpoch AI ↗30 Aug 2026—Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond96.0%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—GPQA Diamond; vendor-estimated API cost/task $0.13436
GSO79.4%Extra highreasoning_effort=xhighIndependent testIndependentGSO ↗27 Sep 2026OpenHandsOpt@1; hack-controlled score 77.45; data https://gso-bench.github.io/assets/leaderboard.json
HealthBench Professional63.4%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—length-adjusted
HLE Diamond60.6%Defaultgpt-6-astraIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 1 (Scale rank accounts for CI); ±3; entry added 2026-09-09
Humanity's Last Exam49.2%LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-low)
Humanity's Last Exam52.7%MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-medium)
Humanity's Last Exam53.1%HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-high)
Humanity's Last Exam54.6%Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-xhigh)
Humanity's Last Exam54.7%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra)
Humanity's Last Exam54.8%DefaultGPT 6 AstraIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 1 (Scale rank accounts for CI); ±1.94; entry added 2026-09-09
Humanity's Last Exam (with tools)57.2%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026—
IOI (Vals)100.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±0 stderr; $6.495/test
LLM Creative Story-Writing Benchmark (Lech Mazur)3.5HighGPT-6 Astra (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 4/56; Thurstone comparison score (centered at 0); est. win chance 91%; 95% bootstrap 3.381 to 3.519
LMArena Code Arena (WebDev)1789Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 2 (CI rank 2-2); 95% CI 1779-1800; 5918 votes
LMArena Text - Coding category1539Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 9 (CI rank 1-36); 95% CI 1526-1552; 2010 votes
LMArena Text - Creative Writing1448Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 39 (rank range 15-77); 95% CI 1433.7-1462.9; 1875 votes; style-controlled
LMArena Text - Expert1513Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 21 (rank range 4-69); 95% CI 1492.5-1533.5; 854 votes; style-controlled
LMArena Text - Hard Prompts1501Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 29 (rank range 11-52); 95% CI 1492.5-1509.6; 5254 votes; style-controlled
LMArena Text - Instruction Following1473Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 30 (rank range 12-58); 95% CI 1461.6-1484.6; 2708 votes; style-controlled
LMArena Text - Longer Query1486Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 34 (rank range 13-60); 95% CI 1475.2-1496.0; 3656 votes; style-controlled
LMArena Text - Multi-Turn1489Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 22 (rank range 6-74); 95% CI 1471.3-1506.1; 1172 votes; style-controlled
LMArena Text - Non-English1462Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 38 (rank range 18-57); 95% CI 1453.5-1470.8; 5098 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1478Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 28 (rank range 6-70); 95% CI 1462.2-1492.8; 1539 votes; style-controlled
LMArena Text - Occupational: Legal & Government1490Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 25 (rank range 1-89); 95% CI 1467.9-1512.4; 728 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1465Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 30 (rank range 11-60); 95% CI 1452.4-1477.9; 2369 votes; style-controlled
OpenAI MRCR v2 (8-needle)100.0%Defaultbest score across efforts (OpenAI table: "maximum at any effort"); 8-needle 256K-512KMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—8-needle, 256K-512K
OpenAI MRCR v2 (8-needle)96.3%Defaultbest score across efforts (OpenAI table: "maximum at any effort"); 8-needle 512K-1MMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—8-needle, 512K-1M
OSWorld 2.0 offline set62.2%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $2.7159
OSWorld 2.0 offline set69.3%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $5.3594
OSWorld 2.0 offline set70.0%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $6.9064
OSWorld 2.0 offline set71.3%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $7.4944
OSWorld 2.0 offline set73.5%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $9.4353
PRBench Finance (Scale)47.5%DefaultGPT 6 AstraIndependent testIndependentScale AI (SEAL) ↗9 Sep 2026—Scale rank 11; ±1.69 CI
PRBench Legal (Scale)48.4%DefaultGPT 6 AstraIndependent testIndependentScale AI (SEAL) ↗9 Sep 2026—Scale rank 10; ±1.64 CI
ProgramBench (fully resolved)5.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.616 stderr; $11.575/test; strict fully-resolved rate
SciCode54.1%LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-low)
SciCode54.2%MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-medium)
SciCode55.4%HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-high)
SciCode55.7%Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-xhigh)
SciCode56.5%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra)
ScreenSpot-Pro91.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.07748
ScreenSpot-Pro91.6%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.07894
ScreenSpot-Pro91.7%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.08518
ScreenSpot-Pro92.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.09311
ScreenSpot-Pro92.6%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026—ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.11097
SimpleBench83.6%DefaultGPT-6 AstraIndependent testIndependentSimpleBench ↗7 Sep 2026—AVG@5, temp 0.7; rank 4th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified75.6%Maxgpt-6-astra_maxIndependent testIndependentEpoch AI ↗30 Aug 2026—Epoch-run (no tools); ±1.36 stderr
SWE Atlas - Codebase QnA59.4%Extra highGPT-6 Astra XHigh + SWE-2 MediumIndependent testIndependentArtificial Analysis ↗1 Oct 2026Devin Fusion CLIAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA59.1%Extra highGPT 6 Astra (Codex) xHigh*Independent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 1 (Scale rank accounts for CI); ±4.88; entry added 2026-09-09; * = Scale footnote (see page)
SWE Atlas - Codebase QnA61.8%MaxGPT-6 Astra (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Refactoring59.0%Extra highGPT 6 Astra (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 1 (Scale rank accounts for CI); ±6.43; entry added 2026-09-09
SWE Atlas - Test Writing50.7%Extra highGPT 6 Astra (Codex) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 1 (Scale rank accounts for CI); ±5.91; entry added 2026-09-09
SWE-Bench Pro V2 (full)96.9%HighGPT-6-Astra (Codex) highIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 4 (Scale rank accounts for CI); ±1.1; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)90.2%HighGPT-6 Astra (Codex) highIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 3 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
Terminal-Bench 2.188.0%LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-low)
Terminal-Bench 2.189.5%MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-medium)
Terminal-Bench 2.189.9%HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-high)
Terminal-Bench 2.189.1%Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-xhigh)
Terminal-Bench 2.188.4%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra)
Terminal-Bench 2.187.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.375 stderr; $1.339/test
Terminal-Bench 2.187.4%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Codex CLITB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 4.050.6%LowIndependent testIndependentTerminal-Bench ↗1 Oct 2026CodexTB 4.0.0 (66 tasks); ±2.75 95% CI; 330 trials; total run cost $1557; model release 2026-09-03
Terminal-Bench 4.041.9%LowGPT-6 Astra (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-low)
Terminal-Bench 4.049.7%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $4.9464
Terminal-Bench 4.054.2%MediumIndependent testIndependentTerminal-Bench ↗1 Oct 2026CodexTB 4.0.0 (66 tasks); ±2.66 95% CI; 330 trials; total run cost $1915; model release 2026-09-03
Terminal-Bench 4.049.5%MediumGPT-6 Astra (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-medium)
Terminal-Bench 4.053.9%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $6.15197
Terminal-Bench 4.057.9%HighIndependent testIndependentTerminal-Bench ↗1 Oct 2026CodexTB 4.0.0 (66 tasks); ±2.97 95% CI; 330 trials; total run cost $2269; model release 2026-09-03
Terminal-Bench 4.054.0%HighGPT-6 Astra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-high)
Terminal-Bench 4.057.9%Highreasoning effort=high (best; max scored lower)Maker's own figureVendor-reportedOpenAI ↗3 Sep 2026Codexheadline Terminal-Bench 4.0 score; at max 56.7%
Terminal-Bench 4.057.9%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $7.20829
Terminal-Bench 4.059.6%Extra highGPT-6 Astra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra-xhigh)
Terminal-Bench 4.057.9%Extra highIndependent testIndependentTerminal-Bench ↗1 Oct 2026CodexTB 4.0.0 (66 tasks); ±2.72 95% CI; 330 trials; total run cost $2351; model release 2026-09-03
Terminal-Bench 4.050.0%Extra highGPT-6 Astra XHigh + SWE-2 MediumIndependent testIndependentArtificial Analysis ↗1 Oct 2026Devin Fusion CLIAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.057.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $7.47814
Terminal-Bench 4.059.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±4.403 stderr; $9.584/test
Terminal-Bench 4.059.1%MaxGPT-6 Astra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-astra)
Terminal-Bench 4.058.2%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026CodexTB 4.0.0 (66 tasks); ±2.79 95% CI; 330 trials; total run cost $3267; model release 2026-09-03
Terminal-Bench 4.055.6%MaxGPT-6 Astra (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.056.7%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗3 Sep 2026CodexTerminal-Bench 4.0; vendor-estimated API cost/task $10.35005
Terminal-Bench-Science 0.155.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $11.4061
Terminal-Bench-Science 0.157.4%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $12.3383
Terminal-Bench-Science 0.162.0%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $14.9534
Terminal-Bench-Science 0.160.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $15.7554
Terminal-Bench-Science 0.168.1%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗29 Sep 2026CodexTerminal-Bench Science 0.1; vendor-estimated API cost/task $23.7974
Vals Code Migration67.7%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.216 stderr; $44.365/test
Vals Index63.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.19 stderr; $18.458/test
Vals Legal Research Bench39.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-6-astra; rank 26/72; ±3.397 stderr; $10.577143/test
Vals SRE Bench56.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±3.066 stderr; $10.238/test
Vals Vibe Code Bench89.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±2.174 stderr; $38.512/test
Vectara Hallucination Leaderboard (HHEM)8.7%Defaultopenai/gpt-6-astraIndependent testIndependentVectara ↗22 Sep 2026—factual consistency 91.3 %; answer rate 100.0 %; avg summary 148.5 words; HHEM-2.3 judge; effort not stated (API default)
Vending-Bench 2$15,515DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$1,074; rank 1; only top 10 rendered server-side (57 more behind 'Show more')