Skip to content
Bencher

Models · OpenAI · Out sinceReleased 9 Jul 2026

GPT-5.6 Luna#23 for coding.#23 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

GPT-5.6 Luna is made by OpenAI. Among the models we track it ranks #23 for coding. It's cheap to use.

Cheap tier (roughly the old "nano" tier). Launched $1/$6, cut 80% on Jul 30, 2026.

Writing & creativity
—
Research & analysis
—
Coding
41.2 / 100 · #23
Price
Cheap$0.20 / $1.20
Price per 1M (blended)Blended / 1M
$0.45
MemoryContext
1.1M
Longest answerMax output
128K
Test resultsResults
90 (44 independent44 indep.)
Out sinceReleased
9 Jul 2026
Made byVendor
OpenAI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-5.6-luna.

Route it as openai/gpt-5.6-luna at $0.20 in / $1.20 out per 1M tokens, 1.1M context. Listed since 9 Jul 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetools

Thinking level

Reasoning effort

How long should GPT-5.6 LunaWhere GPT-5.6 Lunathink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MaxBest at Max

CursorBench

Best at MaxBest at Max

DeepSWE

Best at MaxBest at Max

FrontierCode

Best at MaxBest at Max

Terminal-Bench 2.1

Best at MaxBest at Max

SciCode

Best at MaxBest at Max

Agents' Last Exam

Best at MaxBest at Max

Artificial Analysis Intelligence Index

Best at MaxBest at Max

AutomationBench

Best at MaxBest at Max

OSWorld 2.0 offline set

Best at MaxBest at Max

GPQA Diamond

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Agents' Last Exam31.0%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.1761
Agents' Last Exam38.5%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.375
Agents' Last Exam46.1%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $0.8389
Agents' Last Exam49.4%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $1.5468
Agents' Last Exam50.4%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $2.5698
Agents' Last Exam50.3%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
ARC-AGI-30.2%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
Artificial Analysis Coding Agent Index43.2MaxGPT-5.6 Luna (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 66.4, SWE-Atlas-QnA 48.7, Terminal-Bench v4 14.6; avg cost $0.44/task; avg wall time 24 min/task
Artificial Analysis Coding Agent Index v1.174.6DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—Artificial Analysis Coding Agent Index v1.1 (index score)
Artificial Analysis Intelligence Index34.6Extra highGPT-5.6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-luna-xhigh; list price $0.2/1.2 per 1M in/out; cost to run AA Intelligence Index $0.09/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index37.3MaxGPT-5.6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-5-6-luna; list price $0.2/1.2 per 1M in/out; cost to run AA Intelligence Index $0.18/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index51.2DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—Artificial Analysis Intelligence Index v4.1
Artificial Analysis output speed112 tok/sExtra highGPT-5.6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 52.9s; list price $0.2/1.2 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed109 tok/sMaxGPT-5.6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 110.9s; list price $0.2/1.2 per 1M in/out (AA marks this variant deprecated)
AutomationBench1.8%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.01
AutomationBench4.3%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.02
AutomationBench9.1%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.05
AutomationBench12.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.06
AutomationBench17.0%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.07
AutomationBench14.9%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
BrowseComp83.3%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
CursorBench16.0%LowGPT-5.6 Luna LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 63; $0.03/task; 3,288 tokens/task; 18 steps/task
CursorBench22.2%MediumGPT-5.6 Luna MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 62; $0.08/task; 7,642 tokens/task; 32 steps/task
CursorBench29.4%HighGPT-5.6 Luna HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 52; $0.25/task; 23,368 tokens/task; 64 steps/task
CursorBench33.0%Extra highGPT-5.6 Luna Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 45; $0.44/task; 40,598 tokens/task; 98 steps/task
CursorBench35.9%MaxGPT-5.6 Luna MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 36; $1.03/task; 87,284 tokens/task; 208 steps/task
DeepSWE1.2%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0106
DeepSWE9.3%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.0313
DeepSWE42.4%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.1272
DeepSWE56.9%Extra highgpt-5.6-luna_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 78.8%; ±2.2; 4 runs; $1.54/task
DeepSWE56.2%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.2737
DeepSWE67.2%Maxgpt-5.6-luna_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 90.3%; ±4.0; 4 runs; $3.03/task
DeepSWE66.4%MaxGPT-5.6 Luna (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE62.2%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexDeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $0.532
DeepSWE67.2%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026Codex
Epoch Capabilities Index156.4DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 154.1-158.6; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1829Defaultgpt-5.6-lunaIndependent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 23; rubric score 16.58/20; slop 11.80; avg length 7927 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
EuroEval Swedish (generative)1.6Defaultopenai/gpt-5.6-luna (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 5; ±0.08; lower is better; task scores (first metric): SweDN summarisation 36.55 ± 0.18, Skolprov 85.78 ± 1.83, Swedish facts 71.76 ± 2.30, ScaLA-sv 69.41 ± 0.77
FrontierCode15.4%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.06
FrontierCode25.7%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.13
FrontierCode35.9%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.23
FrontierCode38.9%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.31
FrontierCode39.8%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026CodexFrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $0.37
FrontierCode39.8%Defaultgpt-5.6-luna_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026codexread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness codex; Mean@5
FrontierMath (Tiers 1-3)82.1%Maxgpt-5.6-luna_maxIndependent testIndependentEpoch AI ↗9 Jul 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.3pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 1-378.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
FrontierMath Tier 461.0%Maxgpt-5.6-luna_maxIndependent testIndependentEpoch AI ↗9 Jul 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 7.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierMath Tier 458.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GDP.pdf22.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GDPval-AA v21592DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
GPQA Diamond89.5%Extra highGPT-5.6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated)
GPQA Diamond91.7%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±1.736 stderr; $0.011/test
GPQA Diamond91.1%MaxGPT-5.6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated)
GPQA Diamond92.3%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
HealthBench Professional55.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—official HealthBench Professional scoring
Humanity's Last Exam37.0%Extra highGPT-5.6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated)
Humanity's Last Exam39.5%MaxGPT-5.6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated)
IOI (Vals)61.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±11.613 stderr; $0.580/test
MMMU-Pro (no tools)78.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
MMMU-Pro (with tools)79.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OpenAI MRCR v2 (8-needle)41.3%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 256K-512KMaker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OpenAI MRCR v2 (8-needle)41.3%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results); 8-needle 512K-1MMaker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
OSWorld 2.045.6%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—OSWorld 2.0 (OpenAI-run)
OSWorld 2.0 offline set12.0%Lowreasoning effort=lowMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.0185
OSWorld 2.0 offline set22.7%Mediumreasoning effort=mediumMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.0516
OSWorld 2.0 offline set35.9%Highreasoning effort=highMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.1614
OSWorld 2.0 offline set48.0%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.3573
OSWorld 2.0 offline set52.7%Maxreasoning effort=maxMaker's own figureVendor-reportedOpenAI ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $0.4909
SciCode50.5%Extra highGPT-5.6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated)
SciCode53.6%MaxGPT-5.6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated)
SimpleQA Verified41.0%Maxgpt-5.6-luna_maxIndependent testIndependentEpoch AI ↗10 Aug 2026—Epoch-run (no tools); ±1.56 stderr
SkillsBench60.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.75 stderr; $0.207/test
SWE Atlas - Codebase QnA48.7%MaxGPT-5.6 Luna (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE-Bench Pro (public, v1)62.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
SWE-bench Verified93.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.142 stderr; $0.043/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
SWE-rebench43.6%MediumGPT-5.6 Luna [medium]Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.47; pass@5 59.5%; $0.11/problem
Terminal-Bench 2.177.9%Extra highGPT-5.6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated)
Terminal-Bench 2.180.9%MaxGPT-5.6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated)
Terminal-Bench 2.179.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.991 stderr; $0.054/test
Terminal-Bench 2.175.7%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Codex CLITB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.184.7%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026CodexTerminal-Bench 2.1
Terminal-Bench 3.014.3%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026CodexTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 11.9B, run cost $1.6k
Terminal-Bench 4.03.5%Extra highGPT-5.6 Luna (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna-xhigh) (AA marks this variant deprecated)
Terminal-Bench 4.017.3%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026CodexTB 4.0.0 (66 tasks); ±2.85 95% CI; 330 trials; total run cost $347; model release 2026-06-26
Terminal-Bench 4.014.6%MaxGPT-5.6 Luna (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.011.6%MaxGPT-5.6 Luna (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gpt-5-6-luna) (AA marks this variant deprecated)
Toolathlon53.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗9 Jul 2026—
Vals Code Migration44.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.244 stderr; $1.883/test
Vals Index51.7%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.057 stderr; $0.824/test
Vals TaxEval v276.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id openai/gpt-5.6-luna; rank 7/145; ±0.839 stderr; $0.012217/test