Models · OpenAI · Out sinceReleased 3 Sep 2026
GPT-6 Astra#3 for coding.#3 for coding, best at max effort.
GPT-6 Astra is made by OpenAI. Among the models we track it ranks #3 for coding, #12 for writing, #13 for research and analysis. It's expensive to use. It solves the toughest logic and science puzzles better than any other model.
Flagship. Announced Sep 3, 2026 (limited orgs), API in following days (OpenRouter listing Sep 4). none effort not supported. Cached input $1/M, cache writes $12.5/M. >272K input: 2x input, 1.5x output. Fast mode 2x price; Ultrafast mode available. Meets OpenAI Preparedness "Critical" cyber threshold; misalignment monitoring in production. Knowledge cutoff Apr 30, 2026.
- 57.7 / 100 · #12
- 72.3 / 100 · #13
- 85.4 / 100 · #3
- Expensive$10 / $50
- $20
- 1.1M
- 128K
- 220 (152 independent152 indep.)
- 3 Sep 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-6-astra.
Route it as openai/gpt-6-astra at $10 in / $50 out per 1M tokens, 1.1M context. Listed since 4 Sep 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-6 AstraWhere GPT-6 Astrathink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Artificial Analysis Coding Agent Index
Best at High, worse abovePeaks at High
Terminal-Bench 4.0
Best at Extra high, worse abovePeaks at Extra high
DeepSWE
Best at Extra high, worse abovePeaks at Extra high
FrontierCode
Best at MaxBest at Max
SWE Atlas - Codebase QnA
Best at MaxBest at Max
Terminal-Bench 2.1
Best at High, worse abovePeaks at High
SciCode
Best at MaxBest at Max
Agents' Last Exam
Best at MaxBest at Max
Artificial Analysis Intelligence Index
Best at MaxBest at Max
Terminal-Bench-Science 0.1
Best at MaxBest at Max
AutomationBench
Best at MaxBest at Max
OSWorld 2.0 offline set
Best at MaxBest at Max
AA-LCR
Best at MaxBest at Max
AA-Omniscience Accuracy
Best at MaxBest at Max
AA-Omniscience Hallucination Rate
Best at High, worse abovePeaks at High
Lower is better on this test
AA-Omniscience Index
Best at High, worse abovePeaks at High
ARC-AGI-2
Best at MaxBest at Max
ARC-AGI-3
Best at High, worse abovePeaks at High
FrontierMath Tier 4
Best at MediumBest at Medium
GDP.pdf
Best at Extra high, worse abovePeaks at Extra high
GDPval-AA v2.1
Best at MaxBest at Max
GPQA Diamond
Best at Extra high, worse abovePeaks at Extra high
Humanity's Last Exam
Best at MaxBest at Max
ScreenSpot-Pro
Best at Extra high, worse abovePeaks at Extra high
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 51.3% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| AA-Briefcase v1.1 | 1569 | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 80.0% | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 79.7% | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 80.0% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 80.0% | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 80.7% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 59.5% | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 60.6% | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 61.1% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 61.9% | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 62.6% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 46.9% | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 46.5% | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 44.8% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 48.3% | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 51.3% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 40.6 | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 42.2 | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 43.7 | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 43.4 | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 43.4 | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 53.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $3.0726 |
| Agents' Last Exam | 57.6% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $4.1046 |
| Agents' Last Exam | 57.8% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $4.6396 |
| Agents' Last Exam | 58.3% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $5.397 |
| Agents' Last Exam | 59.3% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $6.2337 |
| ARC-AGI-1 | 98.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| ARC-AGI-2 | 85.4% | LowGPT-6 Astra (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.42/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 92.1% | MediumGPT-6 Astra (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.48/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 92.1% | HighGPT-6 Astra (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.67/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 93.3% | Extra highGPT-6 Astra (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.83/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 95.0% | MaxGPT-6 Astra (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.12/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 95.0% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| ARC-AGI-3 | 96.7% | No reasoningGPT-6 Astra - Provider Adapter (None) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $23,457; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 35.2% | No reasoningGPT-6 Astra (None) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $49,791; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 98.0% | LowGPT-6 Astra - Provider Adapter (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $21,298; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 17.5% | LowGPT-6 Astra (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $38,166; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 98.4% | MediumGPT-6 Astra - Provider Adapter (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $19,285; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 38.6% | MediumGPT-6 Astra (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $48,090; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 99.9% | HighGPT-6 Astra - Provider Adapter (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $18,817; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 54.8% | HighGPT-6 Astra (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $40,705; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 98.4% | Extra highGPT-6 Astra - Provider Adapter (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $18,147; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 59.3% | Extra highGPT-6 Astra (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $37,317; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 98.6% | MaxGPT-6 Astra - Provider Adapter (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $17,332; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 62.7% | MaxGPT-6 Astra (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $26,098; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 99.9% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | run with Responses API harness (two settings changed) |
| Artificial Analysis Coding Agent Index | 62.6 | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $1.50329 |
| Artificial Analysis Coding Agent Index | 65.3 | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $2.32706 |
| Artificial Analysis Coding Agent Index | 65.5 | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.06706 |
| Artificial Analysis Coding Agent Index | 58.9 | Extra highGPT-6 Astra XHigh + SWE-2 Medium | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Devin Fusion CLI | agent Devin Fusion CLI; components: DeepSWE v1.1 67.3, SWE-Atlas-QnA 59.4, Terminal-Bench v4 50.0; avg cost $4.54/task; avg wall time 25 min/task; Devin Fusion = Fable/GPT planner + Cognition SWE-2 (medium) sidekick |
| Artificial Analysis Coding Agent Index | 67 | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.46496 |
| Artificial Analysis Coding Agent Index | 61.6 | MaxGPT-6 Astra (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 67.6, SWE-Atlas-QnA 61.8, Terminal-Bench v4 55.6; avg cost $7.47/task; avg wall time 29 min/task |
| Artificial Analysis Coding Agent Index | 67 | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $4.9813 |
| Artificial Analysis Intelligence Index | 45.8 | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-astra-low; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $0.82/task |
| Artificial Analysis Intelligence Index | 49.6 | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-astra-medium; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $1.54/task |
| Artificial Analysis Intelligence Index | 50.9 | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-astra-high; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $1.73/task |
| Artificial Analysis Intelligence Index | 52.4 | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-astra-xhigh; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $2.31/task |
| Artificial Analysis Intelligence Index | 52.7 | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-astra; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $3.26/task |
| Artificial Analysis Intelligence Index | 61.2 | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | Artificial Analysis Intelligence Index v4.1.1 |
| Artificial Analysis output speed | 43 tok/s | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 2.9s; list price $10/50 per 1M in/out |
| Artificial Analysis output speed | 44 tok/s | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 6.2s; list price $10/50 per 1M in/out |
| Artificial Analysis output speed | 44 tok/s | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 65.6s; list price $10/50 per 1M in/out |
| Artificial Analysis output speed | 48 tok/s | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 205.7s; list price $10/50 per 1M in/out |
| Artificial Analysis output speed | 51 tok/s | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 326.7s; list price $10/50 per 1M in/out |
| AutomationBench | 30.3% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.08 |
| AutomationBench | 34.1% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.27 |
| AutomationBench | 37.1% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.44 |
| AutomationBench | 39.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.5 |
| AutomationBench | 41.4% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $1.73 |
| BenchCAD | 95.9% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | with python tool |
| BrowseComp | 91.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| DeepSWE | 67.0% | Lowgpt-6-astra_low | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 79.6%; ±1.3; 4 runs; $2.19/task |
| DeepSWE | 67.0% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.5952 |
| DeepSWE | 72.8% | Mediumgpt-6-astra_medium | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 82.3%; ±2.6; 4 runs; $4.38/task |
| DeepSWE | 72.8% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.0755 |
| DeepSWE | 73.2% | Highgpt-6-astra_high | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 82.3%; ±3.4; 4 runs; $5.72/task |
| DeepSWE | 73.2% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.9237 |
| DeepSWE | 74.1% | Extra highgpt-6-astra_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 80.5%; ±2.9; 4 runs; $6.52/task |
| DeepSWE | 67.3% | Extra highGPT-6 Astra XHigh + SWE-2 Medium | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Devin Fusion CLI | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 74.1% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $4.4291 |
| DeepSWE | 73.2% | Maxgpt-6-astra_max | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 79.6%; ±0.8; 4 runs; $12.37/task |
| DeepSWE | 67.6% | MaxGPT-6 Astra (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 73.2% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $7.4978 |
| Design Arena (all categories) | 1361 | Maxgpt-6-astra-max | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±9.2 SE; 1521 battles; win rate 60.6% |
| Design Arena (all categories) | 1388 | Defaultgpt-6-astra | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±4.6 SE; 6610 battles; win rate 65.2% |
| Design Arena (fullstack) | 1355 | Defaultgpt-6-astra | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±15.5 SE; 553 battles; win rate 61.1% |
| Epoch Capabilities Index | 166.5 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 163.1-171.1; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 2173 | Default*gpt-6-astra | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 1; rubric score 16.80/20; slop 8.41; avg length 6191 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*) |
| EuroEval Swedish (generative) | 1.2 | Defaultgpt-6-astra (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 1; ±0.04; lower is better; task scores (first metric): SweDN summarisation 39.48 ± 0.17, Skolprov 91.00 ± 1.67, Swedish facts 88.39 ± 1.97, ScaLA-sv 83.01 ± 0.94 |
| FrontierCode | 45.3% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $1.7 |
| FrontierCode | 48.8% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.43 |
| FrontierCode | 50.9% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $3.01 |
| FrontierCode | 50.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $3.28 |
| FrontierCode | 53.3% | Maxgpt-6-astra_max | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | codex | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5 |
| FrontierCode | 53.3% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $4.59 |
| FrontierCode v1.1 (Extended) | 64.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | run with a Codex-like developer message discouraging excess tests/unrelated cleanup |
| FrontierMath (Tiers 1-3) | 93.7% | Maxgpt-6-astra_max | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 82.9% | No reasoninggpt-6-astra_none | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 5.9pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 87.8% | Lowgpt-6-astra_low | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 5.2pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 97.6% | Mediumgpt-6-astra_medium | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 97.6% | Highgpt-6-astra_high | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 97.6% | Extra highgpt-6-astra_xhigh | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 97.6% | Maxgpt-6-astra_max | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 97.6% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | v2 |
| FrontierSWE | 65.5% | Default | Independent testIndependentFrontierSWE ↗ | 1 Oct 2026 | proximus | FrontierSWE V2, mean@5 over 34 tasks (20h budget); ±8.9; $1029.65/trial; 12.1h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort |
| FrontierSWE v2 | 65.5% | Maxreasoning effort=max | Independent testIndependentProximal (via Anthropic system card) ↗ | 22 Sep 2026 | Proximal harness | GPT-6 Astra ranks first on FrontierSWE v2 |
| GDP.pdf | 30.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.69632 |
| GDP.pdf | 30.4% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.71981 |
| GDP.pdf | 31.0% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.79482 |
| GDP.pdf | 32.2% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.91272 |
| GDP.pdf | 31.0% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $2.07582 |
| GDPval-AA v2.1 | 1366 | Lowreasoning effort=low | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $0.85 |
| GDPval-AA v2.1 | 1468 | Mediumreasoning effort=med | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $1.82 |
| GDPval-AA v2.1 | 1485 | Highreasoning effort=high | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $2.43 |
| GDPval-AA v2.1 | 1516 | Extra highreasoning effort=xhigh | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $3.04 |
| GDPval-AA v2.1 | 1542 | Maxreasoning effort=max | Independent testIndependentArtificial Analysis (via Anthropic launch post) ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); cost/task $4.53 |
| GDPval-AA v2.1 | 1542 | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1516.58-1567.2 |
| GPQA Diamond | 93.1% | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-low) |
| GPQA Diamond | 91.8% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.05299 |
| GPQA Diamond | 93.9% | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-medium) |
| GPQA Diamond | 94.9% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.05965 |
| GPQA Diamond | 94.9% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-high) |
| GPQA Diamond | 93.7% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.07805 |
| GPQA Diamond | 96.3% | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-xhigh) |
| GPQA Diamond | 94.4% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.09368 |
| GPQA Diamond | 96.1% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra) |
| GPQA Diamond | 95.8% | Maxgpt-6-astra_max | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 96.0% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | GPQA Diamond; vendor-estimated API cost/task $0.13436 |
| GSO | 79.4% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | 27 Sep 2026 | OpenHands | Opt@1; hack-controlled score 77.45; data https://gso-bench.github.io/assets/leaderboard.json |
| HealthBench Professional | 63.4% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | length-adjusted |
| HLE Diamond | 60.6% | Defaultgpt-6-astra | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 1 (Scale rank accounts for CI); ±3; entry added 2026-09-09 |
| Humanity's Last Exam | 49.2% | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-low) |
| Humanity's Last Exam | 52.7% | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-medium) |
| Humanity's Last Exam | 53.1% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-high) |
| Humanity's Last Exam | 54.6% | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-xhigh) |
| Humanity's Last Exam | 54.7% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra) |
| Humanity's Last Exam | 54.8% | DefaultGPT 6 Astra | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 1 (Scale rank accounts for CI); ±1.94; entry added 2026-09-09 |
| Humanity's Last Exam (with tools) | 57.2% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | |
| IOI (Vals) | 100.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±0 stderr; $6.495/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 3.5 | HighGPT-6 Astra (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 4/56; Thurstone comparison score (centered at 0); est. win chance 91%; 95% bootstrap 3.381 to 3.519 |
| LMArena Code Arena (WebDev) | 1789 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 2 (CI rank 2-2); 95% CI 1779-1800; 5918 votes |
| LMArena Text - Coding category | 1539 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 9 (CI rank 1-36); 95% CI 1526-1552; 2010 votes |
| LMArena Text - Creative Writing | 1448 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 39 (rank range 15-77); 95% CI 1433.7-1462.9; 1875 votes; style-controlled |
| LMArena Text - Expert | 1513 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 21 (rank range 4-69); 95% CI 1492.5-1533.5; 854 votes; style-controlled |
| LMArena Text - Hard Prompts | 1501 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 29 (rank range 11-52); 95% CI 1492.5-1509.6; 5254 votes; style-controlled |
| LMArena Text - Instruction Following | 1473 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 30 (rank range 12-58); 95% CI 1461.6-1484.6; 2708 votes; style-controlled |
| LMArena Text - Longer Query | 1486 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 34 (rank range 13-60); 95% CI 1475.2-1496.0; 3656 votes; style-controlled |
| LMArena Text - Multi-Turn | 1489 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 22 (rank range 6-74); 95% CI 1471.3-1506.1; 1172 votes; style-controlled |
| LMArena Text - Non-English | 1462 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 38 (rank range 18-57); 95% CI 1453.5-1470.8; 5098 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1478 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 28 (rank range 6-70); 95% CI 1462.2-1492.8; 1539 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1490 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 25 (rank range 1-89); 95% CI 1467.9-1512.4; 728 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1465 | Maxgpt-6-astra-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 30 (rank range 11-60); 95% CI 1452.4-1477.9; 2369 votes; style-controlled |
| OpenAI MRCR v2 (8-needle) | 100.0% | Defaultbest score across efforts (OpenAI table: "maximum at any effort"); 8-needle 256K-512K | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | 8-needle, 256K-512K |
| OpenAI MRCR v2 (8-needle) | 96.3% | Defaultbest score across efforts (OpenAI table: "maximum at any effort"); 8-needle 512K-1M | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | 8-needle, 512K-1M |
| OSWorld 2.0 offline set | 62.2% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $2.7159 |
| OSWorld 2.0 offline set | 69.3% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $5.3594 |
| OSWorld 2.0 offline set | 70.0% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $6.9064 |
| OSWorld 2.0 offline set | 71.3% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $7.4944 |
| OSWorld 2.0 offline set | 73.5% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $9.4353 |
| PRBench Finance (Scale) | 47.5% | DefaultGPT 6 Astra | Independent testIndependentScale AI (SEAL) ↗ | 9 Sep 2026 | — | Scale rank 11; ±1.69 CI |
| PRBench Legal (Scale) | 48.4% | DefaultGPT 6 Astra | Independent testIndependentScale AI (SEAL) ↗ | 9 Sep 2026 | — | Scale rank 10; ±1.64 CI |
| ProgramBench (fully resolved) | 5.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.616 stderr; $11.575/test; strict fully-resolved rate |
| SciCode | 54.1% | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-low) |
| SciCode | 54.2% | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-medium) |
| SciCode | 55.4% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-high) |
| SciCode | 55.7% | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-xhigh) |
| SciCode | 56.5% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra) |
| ScreenSpot-Pro | 91.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.07748 |
| ScreenSpot-Pro | 91.6% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.07894 |
| ScreenSpot-Pro | 91.7% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.08518 |
| ScreenSpot-Pro | 92.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.09311 |
| ScreenSpot-Pro | 92.6% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | — | ScreenSpot-Pro, no tools; vendor-estimated API cost/task $0.11097 |
| SimpleBench | 83.6% | DefaultGPT-6 Astra | Independent testIndependentSimpleBench ↗ | 7 Sep 2026 | — | AVG@5, temp 0.7; rank 4th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 75.6% | Maxgpt-6-astra_max | Independent testIndependentEpoch AI ↗ | 30 Aug 2026 | — | Epoch-run (no tools); ±1.36 stderr |
| SWE Atlas - Codebase QnA | 59.4% | Extra highGPT-6 Astra XHigh + SWE-2 Medium | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Devin Fusion CLI | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 59.1% | Extra highGPT 6 Astra (Codex) xHigh* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 1 (Scale rank accounts for CI); ±4.88; entry added 2026-09-09; * = Scale footnote (see page) |
| SWE Atlas - Codebase QnA | 61.8% | MaxGPT-6 Astra (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Refactoring | 59.0% | Extra highGPT 6 Astra (Codex) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 1 (Scale rank accounts for CI); ±6.43; entry added 2026-09-09 |
| SWE Atlas - Test Writing | 50.7% | Extra highGPT 6 Astra (Codex) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 1 (Scale rank accounts for CI); ±5.91; entry added 2026-09-09 |
| SWE-Bench Pro V2 (full) | 96.9% | HighGPT-6-Astra (Codex) high | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 4 (Scale rank accounts for CI); ±1.1; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-Bench Pro V2 (hard) | 90.2% | HighGPT-6 Astra (Codex) high | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 3 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| Terminal-Bench 2.1 | 88.0% | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-low) |
| Terminal-Bench 2.1 | 89.5% | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-medium) |
| Terminal-Bench 2.1 | 89.9% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-high) |
| Terminal-Bench 2.1 | 89.1% | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-xhigh) |
| Terminal-Bench 2.1 | 88.4% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra) |
| Terminal-Bench 2.1 | 87.3% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.375 stderr; $1.339/test |
| Terminal-Bench 2.1 | 87.4% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Codex CLI | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 4.0 | 50.6% | Low | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Codex | TB 4.0.0 (66 tasks); ±2.75 95% CI; 330 trials; total run cost $1557; model release 2026-09-03 |
| Terminal-Bench 4.0 | 41.9% | LowGPT-6 Astra (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-low) |
| Terminal-Bench 4.0 | 49.7% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $4.9464 |
| Terminal-Bench 4.0 | 54.2% | Medium | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Codex | TB 4.0.0 (66 tasks); ±2.66 95% CI; 330 trials; total run cost $1915; model release 2026-09-03 |
| Terminal-Bench 4.0 | 49.5% | MediumGPT-6 Astra (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-medium) |
| Terminal-Bench 4.0 | 53.9% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $6.15197 |
| Terminal-Bench 4.0 | 57.9% | High | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Codex | TB 4.0.0 (66 tasks); ±2.97 95% CI; 330 trials; total run cost $2269; model release 2026-09-03 |
| Terminal-Bench 4.0 | 54.0% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-high) |
| Terminal-Bench 4.0 | 57.9% | Highreasoning effort=high (best; max scored lower) | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | headline Terminal-Bench 4.0 score; at max 56.7% |
| Terminal-Bench 4.0 | 57.9% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $7.20829 |
| Terminal-Bench 4.0 | 59.6% | Extra highGPT-6 Astra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra-xhigh) |
| Terminal-Bench 4.0 | 57.9% | Extra high | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Codex | TB 4.0.0 (66 tasks); ±2.72 95% CI; 330 trials; total run cost $2351; model release 2026-09-03 |
| Terminal-Bench 4.0 | 50.0% | Extra highGPT-6 Astra XHigh + SWE-2 Medium | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Devin Fusion CLI | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 57.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $7.47814 |
| Terminal-Bench 4.0 | 59.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±4.403 stderr; $9.584/test |
| Terminal-Bench 4.0 | 59.1% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-astra) |
| Terminal-Bench 4.0 | 58.2% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Codex | TB 4.0.0 (66 tasks); ±2.79 95% CI; 330 trials; total run cost $3267; model release 2026-09-03 |
| Terminal-Bench 4.0 | 55.6% | MaxGPT-6 Astra (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 56.7% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 3 Sep 2026 | Codex | Terminal-Bench 4.0; vendor-estimated API cost/task $10.35005 |
| Terminal-Bench-Science 0.1 | 55.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $11.4061 |
| Terminal-Bench-Science 0.1 | 57.4% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $12.3383 |
| Terminal-Bench-Science 0.1 | 62.0% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $14.9534 |
| Terminal-Bench-Science 0.1 | 60.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $15.7554 |
| Terminal-Bench-Science 0.1 | 68.1% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 29 Sep 2026 | Codex | Terminal-Bench Science 0.1; vendor-estimated API cost/task $23.7974 |
| Vals Code Migration | 67.7% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.216 stderr; $44.365/test |
| Vals Index | 63.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.19 stderr; $18.458/test |
| Vals Legal Research Bench | 39.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-6-astra; rank 26/72; ±3.397 stderr; $10.577143/test |
| Vals SRE Bench | 56.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±3.066 stderr; $10.238/test |
| Vals Vibe Code Bench | 89.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±2.174 stderr; $38.512/test |
| Vectara Hallucination Leaderboard (HHEM) | 8.7% | Defaultopenai/gpt-6-astra | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 91.3 %; answer rate 100.0 %; avg summary 148.5 words; HHEM-2.3 judge; effort not stated (API default) |
| Vending-Bench 2 | $15,515 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$1,074; rank 1; only top 10 rendered server-side (57 more behind 'Show more') |