Skip to content
Bencher

Models · OpenAI · Out sinceReleased 22 Sep 2026

GPT-6 Luna#28 for coding.#28 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-6 Luna is made by OpenAI. Among the models we track it ranks #28 for coding, #33 for writing, #35 for research and analysis. It's cheap to use.

Current cheap tier. Default effort medium. Cached input $0.01/M. Knowledge cutoff May 18, 2026.

Writing & creativity
13.0 / 100 · #33
Research & analysis
40.1 / 100 · #35
Coding
31.8 / 100 · #28
Price
Cheap$0.10 / $0.50
Price per 1M (blended)Blended / 1M
$0.20
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
76 (51 independent51 indep.)
Out sinceReleased
22 Sep 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-6-luna.

Route it as openai/gpt-6-luna at $0.10 in / $0.50 out per 1M tokens, 1.1M context. Listed since 22 Sep 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-6 LunaWhere GPT-6 Lunathink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MaxBest at Max

DeepSWE

Best at MaxBest at Max

FrontierCode

Best at MaxBest at Max

SciCode

Best at MaxBest at Max

Agents' Last Exam

Best at MaxBest at Max

Artificial Analysis Intelligence Index

Best at MaxBest at Max

AutomationBench

Best at MaxBest at Max

OSWorld 2.0 offline set

Best at MaxBest at Max

AA-LCR

Best at MaxBest at Max

AA-Omniscience Accuracy

Best at Extra high, worse abovePeaks at Extra high

AA-Omniscience Hallucination Rate

Best at MaxBest at Max

Lower is better on this test

AA-Omniscience Index

Best at MaxBest at Max

ARC-AGI-3

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11336MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR80.0%Extra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.3%MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy44.2%Extra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy43.8%MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate82.4%Extra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate76.7%MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index-1.8Extra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index0.7MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam36.3%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.0247
Agents' Last Exam46.8%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1054
Agents' Last Exam43.6%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1117
Agents' Last Exam47.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1105
Agents' Last Exam50.9%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1545
ARC-AGI-30.3%MediumGPT-6 Luna - Provider Adapter (Medium)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $214; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.4%HighGPT-6 Luna - Provider Adapter (High)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $224; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.5%Extra highGPT-6 Luna - Provider Adapter (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $223; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-30.6%MaxGPT-6 Luna - Provider Adapter (Max)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $237; data https://arcprize.org/media/data/leaderboard/v3.json
Artificial Analysis Coding Agent Index41.1MaxGPT-6 Luna (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 63.7, SWE-Atlas-QnA 44.4, Terminal-Bench v4 15.2; avg cost $0.18/task; avg wall time 21 min/task
Artificial Analysis Intelligence Index34.6Extra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-luna-xhigh; list price $0.1/0.5 per 1M in/out; cost to run AA Intelligence Index $0.04/task
Artificial Analysis Intelligence Index38.1MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-6-luna; list price $0.1/0.5 per 1M in/out; cost to run AA Intelligence Index $0.07/task
Artificial Analysis output speed131 tok/sExtra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 23.1s; list price $0.1/0.5 per 1M in/out
Artificial Analysis output speed125 tok/sMaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 123.9s; list price $0.1/0.5 per 1M in/out
AutomationBench1.2%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.006
AutomationBench9.4%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.0163
AutomationBench14.5%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.0208
AutomationBench12.6%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.0245
AutomationBench20.7%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.0367
DeepSWE2.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0057
DeepSWE44.5%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0518
DeepSWE59.3%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0838
DeepSWE61.3%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.1096
DeepSWE63.7%MaxGPT-6 Luna (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE66.6%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.2169
FrontierCode25.7%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.021
FrontierCode35.5%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.0533
FrontierCode37.3%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.0672
FrontierCode37.1%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.0728
FrontierCode42.4%Maxgpt-6-luna_maxIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026codexread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5
FrontierCode42.4%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.1072
FrontierMath (Tiers 1-3)78.9%Maxgpt-6-luna_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 456.1%Maxgpt-6-luna_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 7.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
GDPval-AA v2.11438MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1415.53-1461.24
Humanity's Last Exam34.3%Extra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-luna-xhigh)
Humanity's Last Exam38.5%MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-luna)
IOI (Vals)55.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±8.866 stderr; $0.199/test
LMArena Text - Creative Writing1409Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 89 (rank range 65-137); 95% CI 1392.9-1424.4; 1588 votes; style-controlled
LMArena Text - Expert1498Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 44 (rank range 10-93); 95% CI 1478.7-1517.9; 886 votes; style-controlled
LMArena Text - Hard Prompts1469Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 80 (rank range 57-105); 95% CI 1460.3-1478.1; 4727 votes; style-controlled
LMArena Text - Instruction Following1449Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 65 (rank range 36-90); 95% CI 1437.1-1460.1; 2621 votes; style-controlled
LMArena Text - Longer Query1460Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 73 (rank range 48-103); 95% CI 1449.7-1471.1; 3301 votes; style-controlled
LMArena Text - Multi-Turn1455Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 79 (rank range 33-122); 95% CI 1437.3-1473.4; 1115 votes; style-controlled
LMArena Text - Non-English1436Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 79 (rank range 54-104); 95% CI 1427.3-1445.2; 4677 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1446Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 81 (rank range 36-123); 95% CI 1429.9-1463.0; 1353 votes; style-controlled
LMArena Text - Occupational: Legal & Government1461Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 72 (rank range 18-139); 95% CI 1436.9-1484.5; 618 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1421Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 89 (rank range 67-123); 95% CI 1407.7-1434.6; 2116 votes; style-controlled
OSWorld 2.0 offline set8.3%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.0297
OSWorld 2.0 offline set31.5%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.0624
OSWorld 2.0 offline set41.4%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.1227
OSWorld 2.0 offline set46.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.1639
OSWorld 2.0 offline set52.7%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.2678
PRBench Legal (Scale)42.9%Defaultgpt-6-lunaIndependent testIndependentScale AI (SEAL) ↗25 Sep 2026—Scale rank 18; ±1.69 CI
ProgramBench (fully resolved)0.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.5 stderr; $0.181/test; strict fully-resolved rate
SciCode51.7%Extra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-luna-xhigh)
SciCode54.6%MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-luna)
SimpleQA Verified41.4%Maxgpt-6-luna_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—Epoch-run (no tools); ±1.56 stderr
SWE Atlas - Codebase QnA44.4%MaxGPT-6 Luna (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
Terminal-Bench 2.173.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.716 stderr; $0.029/test
Terminal-Bench 4.08.1%Extra highGPT-6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-luna-xhigh)
Terminal-Bench 4.015.2%MaxGPT-6 Luna (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.012.6%MaxGPT-6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-6-luna)
Vals Code Migration42.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.415 stderr; $0.601/test
Vals Index51.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.072 stderr; $0.431/test
Vals Legal Research Bench30.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-6-luna; rank 40/72; ±3.194 stderr; $0.440296/test
Vals Public Benefits Bench v1.157.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id openai/gpt-6-luna; rank 33/45; ±1.285 stderr; $0.224168/test
Vals Vibe Code Bench81.7%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±3.377 stderr; $1.346/test