Skip to content
Bencher

Models · SpaceXAI (formerly xAI) · Out sinceReleased 12 Aug 2026

Grok 4.6#21 for coding.#21 for coding, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Grok 4.6 is made by SpaceXAI (formerly xAI). Among the models we track it ranks #21 for coding. It's mid-priced to use.

Default reasoning effort = high. >200k-token prompts $4/$12. Fast variant at 2x price.

Writing & creativity
—
Research & analysis
—
Coding
48.9 / 100 · #21
Price
Mid-priced$2 / $6
Price per 1M (blended)Blended / 1M
$3
MemoryContext
500K
Longest answerMax output
—
Test resultsResults
78 (62 independent62 indep.)
Out sinceReleased
12 Aug 2026
Made byVendor
SpaceXAI (formerly xAI)
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra high

Highlighted: where it did best for coding (High thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as x-ai/grok-4.6.

Route it as x-ai/grok-4.6 at $2 in / $6 out per 1M tokens, 500K context. Listed since 12 Aug 2026.

include_reasoninglogprobsmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Thinking level

Reasoning effort

How long should Grok 4.6Where Grok 4.6think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at High, worse abovePeaks at High

CursorBench

Best at Extra highBest at Extra high

DeepSWE

Best at Medium, worse abovePeaks at Medium

Terminal-Bench 2.1

Best at High, worse abovePeaks at High

SciCode

Best at High, worse abovePeaks at High

Artificial Analysis Intelligence Index

Best at High, worse abovePeaks at High

GPQA Diamond

Best at High, worse abovePeaks at High

Humanity's Last Exam

Best at Extra highBest at Extra high

SimpleQA Verified

Best at High, worse abovePeaks at High

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase1577HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—
AA-Briefcase v1.11546HighMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—Comparison column in the Grok 4.7 launch post.
APEX-Agents57.5%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—
APEX-SWE56.4%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—
Artificial Analysis Coding Agent Index47Extra highGrok 4.6 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Grok Buildagent Grok Build; components: DeepSWE v1.1 64.9, SWE-Atlas-QnA 58.3, Terminal-Bench v4 17.7; avg cost $3.57/task; avg wall time 19 min/task
Artificial Analysis Intelligence Index35.1LowGrok 4.6 (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug grok-4-6-low; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $0.48/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index42.8MediumGrok 4.6 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug grok-4-6-medium; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $1.50/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index44.3HighGrok 4.6 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug grok-4-6; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $1.86/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index61HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—
Artificial Analysis Intelligence Index44.2Extra highGrok 4.6 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug grok-4-6-xhigh; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $2.32/task (AA marks this variant deprecated)
Artificial Analysis output speed65 tok/sLowGrok 4.6 (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 7.6s; list price $2/6 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed63 tok/sMediumGrok 4.6 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 27.6s; list price $2/6 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed63 tok/sHighGrok 4.6 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 35.2s; list price $2/6 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed63 tok/sExtra highGrok 4.6 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 31.3s; list price $2/6 per 1M in/out (AA marks this variant deprecated)
CursorBench33.4%LowGrok 4.6 LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 41; $2.25/task; 16,307 tokens/task; 32 steps/task
CursorBench36.1%MediumGrok 4.6 MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 35; $3.48/task; 24,893 tokens/task; 40 steps/task
CursorBench40.4%HighGrok 4.6 HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 27; $5.20/task; 41,387 tokens/task; 48 steps/task
CursorBench40.4%HighMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—Comparison column in the Grok 4.7 launch post.
CursorBench41.4%Extra highGrok 4.6 Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 24; $6.10/task; 49,814 tokens/task; 56 steps/task
CursorBench 3.2.069.9%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—
DeepSWE67.5%Mediumgrok-4.6_mediumIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 84.1%; ±2.3; 4 runs; $3.45/task
DeepSWE65.2%Highgrok-4.6_highIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.0%; ±1.5; 4 runs; $4.38/task
DeepSWE65.2%HighMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—From Grok 4.7 post; Grok 4.6 launch post reported 65.9%. Comparison column in the Grok 4.7 launch post.
DeepSWE66.7%Extra highgrok-4.6_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.0%; ±2.2; 4 runs; $5.50/task
DeepSWE64.9%Extra highGrok 4.6 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Grok BuildAA Coding Agent Index component; pass@1 avg of 3 attempts
EEBench53.0%HighMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—Comparison column in the Grok 4.7 launch post.
Epoch Capabilities Index156.6DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 154.7-158.9; best of listed model versions
FrontierCode48.0%Defaultgrok-4.6_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026grok-buildread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness grok-build; Mean@5
FrontierCode v1.1 (Extended)61.3%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—
GDPval-AA (version unstated)1605HighMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—Elo; version not stated. Comparison column in the Grok 4.7 launch post. Reported by xAI as "GDPval"; Elo scale implies GDPval-AA.
GDPval-AA v21753HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—
GPQA Diamond87.9%LowGrok 4.6 (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated)
GPQA Diamond93.5%MediumGrok 4.6 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated)
GPQA Diamond94.9%HighGrok 4.6 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated)
GPQA Diamond94.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±1.125 stderr; $0.045/test
GPQA Diamond94.0%Highgrok-4.6_highIndependent testIndependentEpoch AI ↗12 Aug 2026—Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond93.5%Extra highGrok 4.6 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated)
GPQA Diamond93.2%Extra highgrok-4.6_xhighIndependent testIndependentEpoch AI ↗14 Aug 2026—Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
Harvey's Legal Agent Benchmark (Vals)15.8%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Harvey LAB (Vals).
HealthBench Professional48.5%HighMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—Comparison column in the Grok 4.7 launch post.
Humanity's Last Exam27.6%LowGrok 4.6 (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated)
Humanity's Last Exam42.1%MediumGrok 4.6 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated)
Humanity's Last Exam42.9%HighGrok 4.6 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated)
Humanity's Last Exam44.1%Extra highGrok 4.6 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated)
LegalBench (Vals)86.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id grok/grok-4.6; rank 13/149; ±0.418 stderr; $0.01457/test
LiveCodeBench88.2%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±0.941 stderr; $0.040/test
LLM Creative Story-Writing Benchmark (Lech Mazur)-2.8HighGrok 4.6 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 51/56; Thurstone comparison score (centered at 0); est. win chance 13%; 95% bootstrap -2.977 to -2.700
LMArena Code Arena (WebDev)1620Highgrok-4.6-highIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 20 (CI rank 15-24); 95% CI 1612-1628; 8120 votes
SciCode49.4%LowGrok 4.6 (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated)
SciCode55.9%MediumGrok 4.6 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated)
SciCode56.5%HighGrok 4.6 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated)
SciCode53.0%Extra highGrok 4.6 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated)
SimpleBench75.9%DefaultGrok 4.6Independent testIndependentSimpleBench ↗13 Aug 2026—AVG@5, temp 0.7; rank 13th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified49.3%Highgrok-4.6_highIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SimpleQA Verified48.9%Extra highgrok-4.6_xhighIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SkillsBench55.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.844 stderr; $0.921/test
SWE Atlas - Codebase QnA58.3%Extra highGrok 4.6 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Grok BuildAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE-bench Verified95.6%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±0.918 stderr; $0.785/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.175.3%LowGrok 4.6 (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated)
Terminal-Bench 2.184.3%MediumGrok 4.6 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated)
Terminal-Bench 2.188.4%HighGrok 4.6 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated)
Terminal-Bench 2.178.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±2.085 stderr; $0.454/test
Terminal-Bench 2.188.0%Extra highGrok 4.6 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated)
Terminal-Bench 3.026.5%HighIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Grok BuildTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 2.9B, run cost $2.1k
Terminal-Bench 3.026.0%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—
Terminal-Bench 4.03.0%LowGrok 4.6 (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated)
Terminal-Bench 4.013.1%MediumGrok 4.6 (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated)
Terminal-Bench 4.021.2%HighGrok 4.6 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated)
Terminal-Bench 4.020.3%HighIndependent testIndependentTerminal-Bench ↗1 Oct 2026Grok BuildTB 4.0.0 (66 tasks); ±3.09 95% CI; 330 trials; total run cost $3592; model release 2026-08-12
Terminal-Bench 4.017.2%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.336 stderr; $5.071/test
Terminal-Bench 4.020.3%HighMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—Comparison column in the Grok 4.7 launch post.
Terminal-Bench 4.017.7%Extra highGrok 4.6 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Grok BuildAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.017.2%Extra highGrok 4.6 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated)
Vals Code Migration44.6%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±4.479 stderr; $14.575/test
Vals Index52.1%Highreasoning_effort=highIndependent testIndependentVals.ai ↗30 Sep 2026—±1.153 stderr; $4.493/test
Vals Legal Research Bench48.1%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id grok/grok-4.6; rank 8/72; ±3.473 stderr; $1.530204/test
Vals Public Benefits Bench v1.166.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id grok/grok-4.6; rank 15/45; ±1.225 stderr; $0.91559/test
Vending-Bench 2$9,047DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$1,604; rank 9; only top 10 rendered server-side (57 more behind 'Show more')