Models · OpenAI · Out sinceReleased 9 Jul 2026
GPT-5.6 Luna#23 for coding.#23 for coding, best at max effort.
GPT-5.6 Luna is made by OpenAI. Among the models we track it ranks #23 for coding. It's cheap to use.
Cheap tier (roughly the old "nano" tier). Launched $1/$6, cut 80% on Jul 30, 2026.
- —
- —
- 41.2 / 100 · #23
- Cheap$0.20 / $1.20
- $0.45
- 1.1M
- 128K
- 90 (44 independent44 indep.)
- 9 Jul 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.6-luna.
Route it as openai/gpt-5.6-luna at $0.20 in / $1.20 out per 1M tokens, 1.1M context. Listed since 9 Jul 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-5.6 LunaWhere GPT-5.6 Lunathink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
DeepSWE
Best at MaxBest at Max
FrontierCode
Best at MaxBest at Max
Terminal-Bench 2.1
Best at MaxBest at Max
SciCode
Best at MaxBest at Max
Agents' Last Exam
Best at MaxBest at Max
Artificial Analysis Intelligence Index
Best at MaxBest at Max
AutomationBench
Best at MaxBest at Max
OSWorld 2.0 offline set
Best at MaxBest at Max
GPQA Diamond
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Agents' Last Exam | 31.0% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1761 |
| Agents' Last Exam | 38.5% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.375 |
| Agents' Last Exam | 46.1% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.8389 |
| Agents' Last Exam | 49.4% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.5468 |
| Agents' Last Exam | 50.4% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $2.5698 |
| Agents' Last Exam | 50.3% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| ARC-AGI-3 | 0.2% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| Artificial Analysis Coding Agent Index | 43.2 | MaxGPT-5.6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 66.4, SWE-Atlas-QnA 48.7, Terminal-Bench v4 14.6; avg cost $0.44/task; avg wall time 24 min/task |
| Artificial Analysis Coding Agent Index v1.1 | 74.6 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | Artificial Analysis Coding Agent Index v1.1 (index score) |
| Artificial Analysis Intelligence Index | 34.6 | Extra highGPT-5.6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-luna-xhigh; list price $0.2/1.2 per 1M in/out; cost to run AA Intelligence Index $0.09/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 37.3 | MaxGPT-5.6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-5-6-luna; list price $0.2/1.2 per 1M in/out; cost to run AA Intelligence Index $0.18/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 51.2 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | Artificial Analysis Intelligence Index v4.1 |
| Artificial Analysis output speed | 112 tok/s | Extra highGPT-5.6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 52.9s; list price $0.2/1.2 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 109 tok/s | MaxGPT-5.6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 110.9s; list price $0.2/1.2 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 1.8% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.01 |
| AutomationBench | 4.3% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.02 |
| AutomationBench | 9.1% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.05 |
| AutomationBench | 12.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.06 |
| AutomationBench | 17.0% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.07 |
| AutomationBench | 14.9% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| BrowseComp | 83.3% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| CursorBench | 16.0% | LowGPT-5.6 Luna Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 63; $0.03/task; 3,288 tokens/task; 18 steps/task |
| CursorBench | 22.2% | MediumGPT-5.6 Luna Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 62; $0.08/task; 7,642 tokens/task; 32 steps/task |
| CursorBench | 29.4% | HighGPT-5.6 Luna High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 52; $0.25/task; 23,368 tokens/task; 64 steps/task |
| CursorBench | 33.0% | Extra highGPT-5.6 Luna Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 45; $0.44/task; 40,598 tokens/task; 98 steps/task |
| CursorBench | 35.9% | MaxGPT-5.6 Luna Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 36; $1.03/task; 87,284 tokens/task; 208 steps/task |
| DeepSWE | 1.2% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0106 |
| DeepSWE | 9.3% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0313 |
| DeepSWE | 42.4% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.1272 |
| DeepSWE | 56.9% | Extra highgpt-5.6-luna_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 78.8%; ±2.2; 4 runs; $1.54/task |
| DeepSWE | 56.2% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.2737 |
| DeepSWE | 67.2% | Maxgpt-5.6-luna_max | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 90.3%; ±4.0; 4 runs; $3.03/task |
| DeepSWE | 66.4% | MaxGPT-5.6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 62.2% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.532 |
| DeepSWE | 67.2% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | |
| Epoch Capabilities Index | 156.4 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 154.1-158.6; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1829 | Defaultgpt-5.6-luna | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 23; rubric score 16.58/20; slop 11.80; avg length 7927 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| EuroEval Swedish (generative) | 1.6 | Defaultopenai/gpt-5.6-luna (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 5; ±0.08; lower is better; task scores (first metric): SweDN summarisation 36.55 ± 0.18, Skolprov 85.78 ± 1.83, Swedish facts 71.76 ± 2.30, ScaLA-sv 69.41 ± 0.77 |
| FrontierCode | 15.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.06 |
| FrontierCode | 25.7% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.13 |
| FrontierCode | 35.9% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.23 |
| FrontierCode | 38.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.31 |
| FrontierCode | 39.8% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.37 |
| FrontierCode | 39.8% | Defaultgpt-5.6-luna_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | codex | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5 |
| FrontierMath (Tiers 1-3) | 82.1% | Maxgpt-5.6-luna_max | Independent testIndependentEpoch AI ↗ | 9 Jul 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.3pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 1-3 | 78.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| FrontierMath Tier 4 | 61.0% | Maxgpt-5.6-luna_max | Independent testIndependentEpoch AI ↗ | 9 Jul 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 7.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierMath Tier 4 | 58.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GDP.pdf | 22.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GDPval-AA v2 | 1592 | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| GPQA Diamond | 89.5% | Extra highGPT-5.6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated) |
| GPQA Diamond | 91.7% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.736 stderr; $0.011/test |
| GPQA Diamond | 91.1% | MaxGPT-5.6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated) |
| GPQA Diamond | 92.3% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| HealthBench Professional | 55.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | official HealthBench Professional scoring |
| Humanity's Last Exam | 37.0% | Extra highGPT-5.6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated) |
| Humanity's Last Exam | 39.5% | MaxGPT-5.6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated) |
| IOI (Vals) | 61.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±11.613 stderr; $0.580/test |
| MMMU-Pro (no tools) | 78.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| MMMU-Pro (with tools) | 79.5% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 41.3% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 256K-512K | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OpenAI MRCR v2 (8-needle) | 41.3% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 512K-1M | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| OSWorld 2.0 | 45.6% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | OSWorld 2.0 (OpenAI-run) |
| OSWorld 2.0 offline set | 12.0% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.0185 |
| OSWorld 2.0 offline set | 22.7% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.0516 |
| OSWorld 2.0 offline set | 35.9% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.1614 |
| OSWorld 2.0 offline set | 48.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.3573 |
| OSWorld 2.0 offline set | 52.7% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.4909 |
| SciCode | 50.5% | Extra highGPT-5.6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated) |
| SciCode | 53.6% | MaxGPT-5.6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated) |
| SimpleQA Verified | 41.0% | Maxgpt-5.6-luna_max | Independent testIndependentEpoch AI ↗ | 10 Aug 2026 | — | Epoch-run (no tools); ±1.56 stderr |
| SkillsBench | 60.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.75 stderr; $0.207/test |
| SWE Atlas - Codebase QnA | 48.7% | MaxGPT-5.6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE-Bench Pro (public, v1) | 62.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| SWE-bench Verified | 93.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.142 stderr; $0.043/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| SWE-rebench | 43.6% | MediumGPT-5.6 Luna [medium] | Independent testIndependentSWE-rebench (Nebius) ↗ | 1 Oct 2026 | SWE-rebench standard scaffold | time window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.47; pass@5 59.5%; $0.11/problem |
| Terminal-Bench 2.1 | 77.9% | Extra highGPT-5.6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 80.9% | MaxGPT-5.6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 79.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.991 stderr; $0.054/test |
| Terminal-Bench 2.1 | 75.7% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Codex CLI | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 84.7% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | Codex | Terminal-Bench 2.1 |
| Terminal-Bench 3.0 | 14.3% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Codex | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 11.9B, run cost $1.6k |
| Terminal-Bench 4.0 | 3.5% | Extra highGPT-5.6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 17.3% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Codex | TB 4.0.0 (66 tasks); ±2.85 95% CI; 330 trials; total run cost $347; model release 2026-06-26 |
| Terminal-Bench 4.0 | 14.6% | MaxGPT-5.6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 11.6% | MaxGPT-5.6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated) |
| Toolathlon | 53.4% | DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results) | Maker's own figureVendor-reportedOpenAI ↗ | 9 Jul 2026 | — | |
| Vals Code Migration | 44.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.244 stderr; $1.883/test |
| Vals Index | 51.7% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.057 stderr; $0.824/test |
| Vals TaxEval v2 | 76.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id openai/gpt-5.6-luna; rank 7/145; ±0.839 stderr; $0.012217/test |