Models · OpenAI · Out sinceReleased 22 Sep 2026
GPT-6 Luna#28 for coding.#28 for coding, best at max effort.
GPT-6 Luna is made by OpenAI. Among the models we track it ranks #28 for coding, #33 for writing, #35 for research and analysis. It's cheap to use.
Current cheap tier. Default effort medium. Cached input $0.01/M. Knowledge cutoff May 18, 2026.
- 13.0 / 100 · #33
- 40.1 / 100 · #35
- 31.8 / 100 · #28
- Cheap$0.10 / $0.50
- $0.20
- 1.1M
- 128K
- 76 (51 independent51 indep.)
- 22 Sep 2026
- OpenAI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-6-luna.
Route it as openai/gpt-6-luna at $0.10 in / $0.50 out per 1M tokens, 1.1M context. Listed since 22 Sep 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools
Thinking level
Reasoning effort
How long should GPT-6 LunaWhere GPT-6 Lunathink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
DeepSWE
Best at MaxBest at Max
FrontierCode
Best at MaxBest at Max
SciCode
Best at MaxBest at Max
Agents' Last Exam
Best at MaxBest at Max
Artificial Analysis Intelligence Index
Best at MaxBest at Max
AutomationBench
Best at MaxBest at Max
OSWorld 2.0 offline set
Best at MaxBest at Max
AA-LCR
Best at MaxBest at Max
AA-Omniscience Accuracy
Best at Extra high, worse abovePeaks at Extra high
AA-Omniscience Hallucination Rate
Best at MaxBest at Max
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
ARC-AGI-3
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1336 | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 80.0% | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.3% | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 44.2% | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 43.8% | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 82.4% | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 76.7% | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | -1.8 | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 0.7 | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 36.3% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.0247 |
| Agents' Last Exam | 46.8% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1054 |
| Agents' Last Exam | 43.6% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1117 |
| Agents' Last Exam | 47.9% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1105 |
| Agents' Last Exam | 50.9% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1545 |
| ARC-AGI-3 | 0.3% | MediumGPT-6 Luna - Provider Adapter (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $214; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.4% | HighGPT-6 Luna - Provider Adapter (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $224; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.5% | Extra highGPT-6 Luna - Provider Adapter (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $223; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 0.6% | MaxGPT-6 Luna - Provider Adapter (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $237; data https://arcprize.org/media/data/leaderboard/v3.json |
| Artificial Analysis Coding Agent Index | 41.1 | MaxGPT-6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 63.7, SWE-Atlas-QnA 44.4, Terminal-Bench v4 15.2; avg cost $0.18/task; avg wall time 21 min/task |
| Artificial Analysis Intelligence Index | 34.6 | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-luna-xhigh; list price $0.1/0.5 per 1M in/out; cost to run AA Intelligence Index $0.04/task |
| Artificial Analysis Intelligence Index | 38.1 | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-6-luna; list price $0.1/0.5 per 1M in/out; cost to run AA Intelligence Index $0.07/task |
| Artificial Analysis output speed | 131 tok/s | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 23.1s; list price $0.1/0.5 per 1M in/out |
| Artificial Analysis output speed | 125 tok/s | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 123.9s; list price $0.1/0.5 per 1M in/out |
| AutomationBench | 1.2% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.006 |
| AutomationBench | 9.4% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.0163 |
| AutomationBench | 14.5% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.0208 |
| AutomationBench | 12.6% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.0245 |
| AutomationBench | 20.7% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.0367 |
| DeepSWE | 2.4% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0057 |
| DeepSWE | 44.5% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0518 |
| DeepSWE | 59.3% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0838 |
| DeepSWE | 61.3% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.1096 |
| DeepSWE | 63.7% | MaxGPT-6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 66.6% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.2169 |
| FrontierCode | 25.7% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.021 |
| FrontierCode | 35.5% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.0533 |
| FrontierCode | 37.3% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.0672 |
| FrontierCode | 37.1% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.0728 |
| FrontierCode | 42.4% | Maxgpt-6-luna_max | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | codex | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5 |
| FrontierCode | 42.4% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | Codex | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.1072 |
| FrontierMath (Tiers 1-3) | 78.9% | Maxgpt-6-luna_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 56.1% | Maxgpt-6-luna_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 7.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| GDPval-AA v2.1 | 1438 | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1415.53-1461.24 |
| Humanity's Last Exam | 34.3% | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-luna-xhigh) |
| Humanity's Last Exam | 38.5% | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-luna) |
| IOI (Vals) | 55.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±8.866 stderr; $0.199/test |
| LMArena Text - Creative Writing | 1409 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 89 (rank range 65-137); 95% CI 1392.9-1424.4; 1588 votes; style-controlled |
| LMArena Text - Expert | 1498 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 44 (rank range 10-93); 95% CI 1478.7-1517.9; 886 votes; style-controlled |
| LMArena Text - Hard Prompts | 1469 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 80 (rank range 57-105); 95% CI 1460.3-1478.1; 4727 votes; style-controlled |
| LMArena Text - Instruction Following | 1449 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 65 (rank range 36-90); 95% CI 1437.1-1460.1; 2621 votes; style-controlled |
| LMArena Text - Longer Query | 1460 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 73 (rank range 48-103); 95% CI 1449.7-1471.1; 3301 votes; style-controlled |
| LMArena Text - Multi-Turn | 1455 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 79 (rank range 33-122); 95% CI 1437.3-1473.4; 1115 votes; style-controlled |
| LMArena Text - Non-English | 1436 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 79 (rank range 54-104); 95% CI 1427.3-1445.2; 4677 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1446 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 81 (rank range 36-123); 95% CI 1429.9-1463.0; 1353 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1461 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 72 (rank range 18-139); 95% CI 1436.9-1484.5; 618 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1421 | Maxgpt-6-luna-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 89 (rank range 67-123); 95% CI 1407.7-1434.6; 2116 votes; style-controlled |
| OSWorld 2.0 offline set | 8.3% | Lowreasoning effort=low | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.0297 |
| OSWorld 2.0 offline set | 31.5% | Mediumreasoning effort=medium | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.0624 |
| OSWorld 2.0 offline set | 41.4% | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.1227 |
| OSWorld 2.0 offline set | 46.7% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.1639 |
| OSWorld 2.0 offline set | 52.7% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.2678 |
| PRBench Legal (Scale) | 42.9% | Defaultgpt-6-luna | Independent testIndependentScale AI (SEAL) ↗ | 25 Sep 2026 | — | Scale rank 18; ±1.69 CI |
| ProgramBench (fully resolved) | 0.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.5 stderr; $0.181/test; strict fully-resolved rate |
| SciCode | 51.7% | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-luna-xhigh) |
| SciCode | 54.6% | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-luna) |
| SimpleQA Verified | 41.4% | Maxgpt-6-luna_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | Epoch-run (no tools); ±1.56 stderr |
| SWE Atlas - Codebase QnA | 44.4% | MaxGPT-6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| Terminal-Bench 2.1 | 73.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.716 stderr; $0.029/test |
| Terminal-Bench 4.0 | 8.1% | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-luna-xhigh) |
| Terminal-Bench 4.0 | 15.2% | MaxGPT-6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 12.6% | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gpt-6-luna) |
| Vals Code Migration | 42.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.415 stderr; $0.601/test |
| Vals Index | 51.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.072 stderr; $0.431/test |
| Vals Legal Research Bench | 30.3% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-6-luna; rank 40/72; ±3.194 stderr; $0.440296/test |
| Vals Public Benefits Bench v1.1 | 57.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id openai/gpt-6-luna; rank 33/45; ±1.285 stderr; $0.224168/test |
| Vals Vibe Code Bench | 81.7% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±3.377 stderr; $1.346/test |