Skip to content
Bencher

Models · OpenAI · Out sinceReleased 9 Jul 2026

GPT-5.6 Terra#14 for coding.#14 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-5.6 Terra is made by OpenAI. Among the models we track it ranks #14 for coding, #27 for research and analysis. It's mid-priced to use.

Mid tier (roughly the old "mini" tier). Launched $2.50/$15, cut 20% on Jul 30, 2026.

Writing & creativity
—
Research & analysis
58.7 / 100 · #27
Coding
53.2 / 100 · #14
Price
Mid-priced$2 / $12
Price per 1M (blended)Blended / 1M
$4.50
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
91 (70 independent70 indep.)
Out sinceReleased
9 Jul 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.6-terra.

Route it as openai/gpt-5.6-terra at $2 in / $12 out per 1M tokens, 1.1M context. Listed since 9 Jul 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-5.6 TerraWhere GPT-5.6 Terrathink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MaxBest at Max

CursorBench

Best at MaxBest at Max

DeepSWE

Best at MaxBest at Max

Terminal-Bench 2.1

Best at MaxBest at Max

SciCode

Best at MaxBest at Max

Artificial Analysis Intelligence Index

Best at MaxBest at Max

tau2-bench

Best at MaxBest at Max

AA-LCR

Best at MaxBest at Max

AA-Omniscience Accuracy

Best at MaxBest at Max

AA-Omniscience Hallucination Rate

Best at MaxBest at Max

Lower is better on this test

AA-Omniscience Index

Best at MaxBest at Max

ARC-AGI-3

Best at MaxBest at Max

GPQA Diamond

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR77.7%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR79.0%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.0%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy45.5%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy45.5%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy46.8%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate89.8%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate89.0%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate87.9%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index-3.5HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index-3Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index0.1MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam50.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
ARC-AGI-283.9%MaxGPT-5.6 Terra (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.09/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-30.5%HighGPT-5.6 Terra (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $5,881; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.7%Extra highGPT-5.6 Terra (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $6,804; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.8%MaxGPT-5.6 Terra (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $7,918; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.8%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
Artificial Analysis Coding Agent Index v1.177.4DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—Artificial Analysis Coding Agent Index v1.1 (index score)
Artificial Analysis Intelligence Index34.2HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-terra-high; list price $2/12 per 1M in/out; cost to run AA Intelligence Index $0.34/task
Artificial Analysis Intelligence Index38Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-terra-xhigh; list price $2/12 per 1M in/out; cost to run AA Intelligence Index $0.63/task
Artificial Analysis Intelligence Index42.1MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-terra; list price $2/12 per 1M in/out; cost to run AA Intelligence Index $1.40/task
Artificial Analysis Intelligence Index55DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—Artificial Analysis Intelligence Index v4.1
Artificial Analysis output speed72 tok/sHighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 3.0s; list price $2/12 per 1M in/out
Artificial Analysis output speed80 tok/sExtra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 45.5s; list price $2/12 per 1M in/out
Artificial Analysis output speed89 tok/sMaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 185.5s; list price $2/12 per 1M in/out
AutomationBench15.2%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
BrowseComp87.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
CursorBench25.2%LowGPT-5.6 Terra LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 58; $0.52/task; 5,914 tokens/task; 23 steps/task
CursorBench27.6%MediumGPT-5.6 Terra MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 56; $0.64/task; 7,307 tokens/task; 25 steps/task
CursorBench30.7%HighGPT-5.6 Terra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 51; $1.11/task; 13,162 tokens/task; 33 steps/task
CursorBench33.6%Extra highGPT-5.6 Terra Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 40; $1.81/task; 23,436 tokens/task; 43 steps/task
CursorBench41.3%MaxGPT-5.6 Terra MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 25; $5.14/task; 60,814 tokens/task; 107 steps/task
DeepSWE60.2%Extra highgpt-5.6-terra_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 80.5%; ±2.1; 4 runs; $2.13/task
DeepSWE69.6%Maxgpt-5.6-terra_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 88.5%; ±2.6; 4 runs; $4.95/task
DeepSWE69.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026Codex
Epoch Capabilities Index159.8DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 157.1-162.8; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1855Defaultgpt-5.6-terraIndependent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 17; rubric score 16.56/20; slop 12.40; avg length 10271 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
EuroEval Swedish (generative)1.5Defaultgpt-5.6-terra (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 3; ±0.07; lower is better; task scores (first metric): SweDN summarisation 37.83 ± 0.18, Skolprov 85.93 ± 1.89, Swedish facts 75.78 ± 2.13, ScaLA-sv 71.42 ± 1.14
FrontierCode41.3%Defaultgpt-5.6-terra_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026codexread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5
FrontierMath (Tiers 1-3)86.0%Maxgpt-5.6-terra_maxIndependent testIndependentEpoch AI ↗9 Jul 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 1-384.9%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
FrontierMath Tier 470.7%Maxgpt-5.6-terra_maxIndependent testIndependentEpoch AI ↗9 Jul 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 7.2pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 468.3%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GDP.pdf24.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GDPval-AA v21593DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GPQA Diamond89.6%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-high)
GPQA Diamond90.9%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026—±1.906 stderr; $0.034/test
GPQA Diamond90.8%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-xhigh)
GPQA Diamond93.3%Maxgpt-5.6-terra_maxIndependent testIndependentEpoch AI ↗9 Jul 2026—Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond92.5%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra)
GPQA Diamond92.9%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
HealthBench Professional57.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—official HealthBench Professional scoring
Humanity's Last Exam38.5%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-high)
Humanity's Last Exam41.9%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-xhigh)
Humanity's Last Exam42.9%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra)
IOI (Vals)87.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±6.531 stderr; $8.664/test
MMMU-Pro (no tools)80.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
MMMU-Pro (with tools)82.0%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OpenAI MRCR v2 (8-needle)89.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 256K-512KMaker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OpenAI MRCR v2 (8-needle)72.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 512K-1MMaker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OSWorld 2.050.2%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—OSWorld 2.0 (OpenAI-run)
ProgramBench (fully resolved)0.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.5 stderr; $5.378/test; strict fully-resolved rate
SciCode52.4%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-high)
SciCode52.3%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-xhigh)
SciCode55.0%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra)
SimpleQA Verified43.2%Maxgpt-5.6-terra_maxIndependent testIndependentEpoch AI ↗10 Aug 2026—Epoch-run (no tools); ±1.57 stderr
SkillsBench58.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.472 stderr; $1.752/test
SWE-Bench Pro (public, v1)63.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
SWE-Bench Pro V2 (full)92.4%Extra highGPT-5.6-Terra (Codex) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 9 (Scale rank accounts for CI); ±1.81; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)86.3%Extra highGPT-5.6 Terra (Codex) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Codexrank 6 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-bench Verified95.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±0.938 stderr; $0.401/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
tau2-bench78.4%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-high)
tau2-bench80.4%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-xhigh)
tau2-bench86.3%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra)
Terminal-Bench 2.175.7%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-high)
Terminal-Bench 2.180.1%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-xhigh)
Terminal-Bench 2.188.0%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra)
Terminal-Bench 2.177.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±2.247 stderr; $0.473/test
Terminal-Bench 2.178.4%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Codex CLITB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.187.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026CodexTerminal-Bench 2.1
Terminal-Bench 3.020.8%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026CodexTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 7.0B, run cost $2.5k
Terminal-Bench 4.01.5%HighGPT-5.6 Terra (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-high)
Terminal-Bench 4.010.1%Extra highGPT-5.6 Terra (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra-xhigh)
Terminal-Bench 4.035.4%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-terra)
Terminal-Bench 4.022.7%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±2.314 stderr; $5.602/test
Terminal-Bench 4.021.5%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026CodexTB 4.0.0 (66 tasks); ±3.25 95% CI; 330 trials; total run cost $1734; model release 2026-06-26
Toolathlon53.1%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
Vals Code Migration47.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.28 stderr; $8.127/test
Vals Index53.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.293 stderr; $5.776/test
Vals TaxEval v276.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026—vals id openai/gpt-5.6-terra; rank 6/145; ±0.836 stderr; $0.035909/test