Models · OpenAI · Out sinceReleased 9 Jul 2026
GPT-5.6 Terra#14 for coding.#14 for coding, best at max effort.
GPT-5.6 Terra is made by OpenAI. Among the models we track it ranks #14 for coding, #27 for research and analysis. It's mid-priced to use.
Mid tier (roughly the old "mini" tier). Launched $2.50/$15, cut 20% on Jul 30, 2026.
- —
- 58.7 / 100 · #27
- 53.2 / 100 · #14
- Mid-priced$2 / $12
- $4.50
- 1.1M
- 128K
- 91 (70 independent70 indep.)
- 9 Jul 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.6-terra.
Route it as openai/gpt-5.6-terra at $2 in / $12 out per 1M tokens, 1.1M context. Listed since 9 Jul 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-5.6 TerraWhere GPT-5.6 Terrathink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
DeepSWE
Best at MaxBest at Max
Terminal-Bench 2.1
Best at MaxBest at Max
SciCode
Best at MaxBest at Max
Artificial Analysis Intelligence Index
Best at MaxBest at Max
tau2-bench
Best at MaxBest at Max
AA-LCR
Best at MaxBest at Max
AA-Omniscience Accuracy
Best at MaxBest at Max
AA-Omniscience Hallucination Rate
Best at MaxBest at Max
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
ARC-AGI-3
Best at MaxBest at Max
GPQA Diamond
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-LCR | 77.7% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 79.0% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.0% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 45.5% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 45.5% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 46.8% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 89.8% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 89.0% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 87.9% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | -3.5 | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | -3 | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 0.1 | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 50.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| ARC-AGI-2 | 83.9% | MaxGPT-5.6 Terra (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.09/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-3 | 0.5% | HighGPT-5.6 Terra (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $5,881; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.7% | Extra highGPT-5.6 Terra (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $6,804; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.8% | MaxGPT-5.6 Terra (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $7,918; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.8% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| Artificial Analysis Coding Agent Index v1.1 | 77.4 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | Artificial Analysis Coding Agent Index v1.1 (index score) |
| Artificial Analysis Intelligence Index | 34.2 | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-terra-high; list price $2/12 per 1M in/out; cost to run AA Intelligence Index $0.34/task |
| Artificial Analysis Intelligence Index | 38 | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-terra-xhigh; list price $2/12 per 1M in/out; cost to run AA Intelligence Index $0.63/task |
| Artificial Analysis Intelligence Index | 42.1 | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-terra; list price $2/12 per 1M in/out; cost to run AA Intelligence Index $1.40/task |
| Artificial Analysis Intelligence Index | 55 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | Artificial Analysis Intelligence Index v4.1 |
| Artificial Analysis output speed | 72 tok/s | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 3.0s; list price $2/12 per 1M in/out |
| Artificial Analysis output speed | 80 tok/s | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 45.5s; list price $2/12 per 1M in/out |
| Artificial Analysis output speed | 89 tok/s | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 185.5s; list price $2/12 per 1M in/out |
| AutomationBench | 15.2% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| BrowseComp | 87.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| CursorBench | 25.2% | LowGPT-5.6 Terra Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 58; $0.52/task; 5,914 tokens/task; 23 steps/task |
| CursorBench | 27.6% | MediumGPT-5.6 Terra Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 56; $0.64/task; 7,307 tokens/task; 25 steps/task |
| CursorBench | 30.7% | HighGPT-5.6 Terra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 51; $1.11/task; 13,162 tokens/task; 33 steps/task |
| CursorBench | 33.6% | Extra highGPT-5.6 Terra Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 40; $1.81/task; 23,436 tokens/task; 43 steps/task |
| CursorBench | 41.3% | MaxGPT-5.6 Terra Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 25; $5.14/task; 60,814 tokens/task; 107 steps/task |
| DeepSWE | 60.2% | Extra highgpt-5.6-terra_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 80.5%; ±2.1; 4 runs; $2.13/task |
| DeepSWE | 69.6% | Maxgpt-5.6-terra_max | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 88.5%; ±2.6; 4 runs; $4.95/task |
| DeepSWE | 69.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | |
| Epoch Capabilities Index | 159.8 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 157.1-162.8; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1855 | Defaultgpt-5.6-terra | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 17; rubric score 16.56/20; slop 12.40; avg length 10271 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| EuroEval Swedish (generative) | 1.5 | Defaultgpt-5.6-terra (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 3; ±0.07; lower is better; task scores (first metric): SweDN summarisation 37.83 ± 0.18, Skolprov 85.93 ± 1.89, Swedish facts 75.78 ± 2.13, ScaLA-sv 71.42 ± 1.14 |
| FrontierCode | 41.3% | Defaultgpt-5.6-terra_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | codex | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5 |
| FrontierMath (Tiers 1-3) | 86.0% | Maxgpt-5.6-terra_max | Independent testIndependentEpoch AI ↗ | 9 Jul 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 1-3 | 84.9% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| FrontierMath Tier 4 | 70.7% | Maxgpt-5.6-terra_max | Independent testIndependentEpoch AI ↗ | 9 Jul 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 7.2pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 68.3% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GDP.pdf | 24.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GDPval-AA v2 | 1593 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GPQA Diamond | 89.6% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-high) |
| GPQA Diamond | 90.9% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.906 stderr; $0.034/test |
| GPQA Diamond | 90.8% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-xhigh) |
| GPQA Diamond | 93.3% | Maxgpt-5.6-terra_max | Independent testIndependentEpoch AI ↗ | 9 Jul 2026 | — | Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 92.5% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra) |
| GPQA Diamond | 92.9% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| HealthBench Professional | 57.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | official HealthBench Professional scoring |
| Humanity's Last Exam | 38.5% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-high) |
| Humanity's Last Exam | 41.9% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-xhigh) |
| Humanity's Last Exam | 42.9% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra) |
| IOI (Vals) | 87.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±6.531 stderr; $8.664/test |
| MMMU-Pro (no tools) | 80.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| MMMU-Pro (with tools) | 82.0% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 89.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 256K-512K | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 72.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 512K-1M | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OSWorld 2.0 | 50.2% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | OSWorld 2.0 (OpenAI-run) |
| ProgramBench (fully resolved) | 0.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.5 stderr; $5.378/test; strict fully-resolved rate |
| SciCode | 52.4% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-high) |
| SciCode | 52.3% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-xhigh) |
| SciCode | 55.0% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra) |
| SimpleQA Verified | 43.2% | Maxgpt-5.6-terra_max | Independent testIndependentEpoch AI ↗ | 10 Aug 2026 | — | Epoch-run (no tools); ±1.57 stderr |
| SkillsBench | 58.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.472 stderr; $1.752/test |
| SWE-Bench Pro (public, v1) | 63.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| SWE-Bench Pro V2 (full) | 92.4% | Extra highGPT-5.6-Terra (Codex) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 9 (Scale rank accounts for CI); ±1.81; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-Bench Pro V2 (hard) | 86.3% | Extra highGPT-5.6 Terra (Codex) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Codex | rank 6 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-bench Verified | 95.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±0.938 stderr; $0.401/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| tau2-bench | 78.4% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-high) |
| tau2-bench | 80.4% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-xhigh) |
| tau2-bench | 86.3% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra) |
| Terminal-Bench 2.1 | 75.7% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-high) |
| Terminal-Bench 2.1 | 80.1% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-xhigh) |
| Terminal-Bench 2.1 | 88.0% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra) |
| Terminal-Bench 2.1 | 77.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±2.247 stderr; $0.473/test |
| Terminal-Bench 2.1 | 78.4% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Codex CLI | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 87.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | Terminal-Bench 2.1 |
| Terminal-Bench 3.0 | 20.8% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Codex | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 7.0B, run cost $2.5k |
| Terminal-Bench 4.0 | 1.5% | HighGPT-5.6 Terra (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-high) |
| Terminal-Bench 4.0 | 10.1% | Extra highGPT-5.6 Terra (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra-xhigh) |
| Terminal-Bench 4.0 | 35.4% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-terra) |
| Terminal-Bench 4.0 | 22.7% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±2.314 stderr; $5.602/test |
| Terminal-Bench 4.0 | 21.5% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Codex | TB 4.0.0 (66 tasks); ±3.25 95% CI; 330 trials; total run cost $1734; model release 2026-06-26 |
| Toolathlon | 53.1% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| Vals Code Migration | 47.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.28 stderr; $8.127/test |
| Vals Index | 53.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.293 stderr; $5.776/test |
| Vals TaxEval v2 | 76.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id openai/gpt-5.6-terra; rank 6/145; ±0.836 stderr; $0.035909/test |