Models · SpaceXAI (formerly xAI) · Out sinceReleased 12 Aug 2026
Grok 4.6#21 for coding.#21 for coding, best at high effort.
Grok 4.6 is made by SpaceXAI (formerly xAI). Among the models we track it ranks #21 for coding. It's mid-priced to use.
Default reasoning effort = high. >200k-token prompts $4/$12. Fast variant at 2x price.
- —
- —
- 48.9 / 100 · #21
- Mid-priced$2 / $6
- $3
- 500K
- —
- 78 (62 independent62 indep.)
- 12 Aug 2026
- SpaceXAI (formerly xAI)
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (High thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as x-ai/grok-4.6.
Route it as x-ai/grok-4.6 at $2 in / $6 out per 1M tokens, 500K context. Listed since 12 Aug 2026.
include_reasoninglogprobsmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Thinking level
Reasoning effort
How long should Grok 4.6Where Grok 4.6think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at High, worse abovePeaks at High
CursorBench
Best at Extra highBest at Extra high
DeepSWE
Best at Medium, worse abovePeaks at Medium
Terminal-Bench 2.1
Best at High, worse abovePeaks at High
SciCode
Best at High, worse abovePeaks at High
Artificial Analysis Intelligence Index
Best at High, worse abovePeaks at High
GPQA Diamond
Best at High, worse abovePeaks at High
Humanity's Last Exam
Best at Extra highBest at Extra high
SimpleQA Verified
Best at High, worse abovePeaks at High
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase | 1577 | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | |
| AA-Briefcase v1.1 | 1546 | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | Comparison column in the Grok 4.7 launch post. |
| APEX-Agents | 57.5% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | |
| APEX-SWE | 56.4% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | |
| Artificial Analysis Coding Agent Index | 47 | Extra highGrok 4.6 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Grok Build | agent Grok Build; components: DeepSWE v1.1 64.9, SWE-Atlas-QnA 58.3, Terminal-Bench v4 17.7; avg cost $3.57/task; avg wall time 19 min/task |
| Artificial Analysis Intelligence Index | 35.1 | LowGrok 4.6 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug grok-4-6-low; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $0.48/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 42.8 | MediumGrok 4.6 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug grok-4-6-medium; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $1.50/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 44.3 | HighGrok 4.6 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug grok-4-6; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $1.86/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 61 | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | |
| Artificial Analysis Intelligence Index | 44.2 | Extra highGrok 4.6 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug grok-4-6-xhigh; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $2.32/task (AA marks this variant deprecated) |
| Artificial Analysis output speed | 65 tok/s | LowGrok 4.6 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 7.6s; list price $2/6 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 63 tok/s | MediumGrok 4.6 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 27.6s; list price $2/6 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 63 tok/s | HighGrok 4.6 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 35.2s; list price $2/6 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 63 tok/s | Extra highGrok 4.6 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 31.3s; list price $2/6 per 1M in/out (AA marks this variant deprecated) |
| CursorBench | 33.4% | LowGrok 4.6 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 41; $2.25/task; 16,307 tokens/task; 32 steps/task |
| CursorBench | 36.1% | MediumGrok 4.6 Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 35; $3.48/task; 24,893 tokens/task; 40 steps/task |
| CursorBench | 40.4% | HighGrok 4.6 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 27; $5.20/task; 41,387 tokens/task; 48 steps/task |
| CursorBench | 40.4% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | Comparison column in the Grok 4.7 launch post. |
| CursorBench | 41.4% | Extra highGrok 4.6 Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 24; $6.10/task; 49,814 tokens/task; 56 steps/task |
| CursorBench 3.2.0 | 69.9% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | |
| DeepSWE | 67.5% | Mediumgrok-4.6_medium | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 84.1%; ±2.3; 4 runs; $3.45/task |
| DeepSWE | 65.2% | Highgrok-4.6_high | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.0%; ±1.5; 4 runs; $4.38/task |
| DeepSWE | 65.2% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | From Grok 4.7 post; Grok 4.6 launch post reported 65.9%. Comparison column in the Grok 4.7 launch post. |
| DeepSWE | 66.7% | Extra highgrok-4.6_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.0%; ±2.2; 4 runs; $5.50/task |
| DeepSWE | 64.9% | Extra highGrok 4.6 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Grok Build | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| EEBench | 53.0% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | Comparison column in the Grok 4.7 launch post. |
| Epoch Capabilities Index | 156.6 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 154.7-158.9; best of listed model versions |
| FrontierCode | 48.0% | Defaultgrok-4.6_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | grok-build | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness grok-build; Mean@5 |
| FrontierCode v1.1 (Extended) | 61.3% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | |
| GDPval-AA (version unstated) | 1605 | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | Elo; version not stated. Comparison column in the Grok 4.7 launch post. Reported by xAI as "GDPval"; Elo scale implies GDPval-AA. |
| GDPval-AA v2 | 1753 | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | |
| GPQA Diamond | 87.9% | LowGrok 4.6 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated) |
| GPQA Diamond | 93.5% | MediumGrok 4.6 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated) |
| GPQA Diamond | 94.9% | HighGrok 4.6 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated) |
| GPQA Diamond | 94.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.125 stderr; $0.045/test |
| GPQA Diamond | 94.0% | Highgrok-4.6_high | Independent testIndependentEpoch AI ↗ | 12 Aug 2026 | — | Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 93.5% | Extra highGrok 4.6 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated) |
| GPQA Diamond | 93.2% | Extra highgrok-4.6_xhigh | Independent testIndependentEpoch AI ↗ | 14 Aug 2026 | — | Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| Harvey's Legal Agent Benchmark (Vals) | 15.8% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Harvey LAB (Vals). |
| HealthBench Professional | 48.5% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | Comparison column in the Grok 4.7 launch post. |
| Humanity's Last Exam | 27.6% | LowGrok 4.6 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated) |
| Humanity's Last Exam | 42.1% | MediumGrok 4.6 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated) |
| Humanity's Last Exam | 42.9% | HighGrok 4.6 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated) |
| Humanity's Last Exam | 44.1% | Extra highGrok 4.6 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated) |
| LegalBench (Vals) | 86.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id grok/grok-4.6; rank 13/149; ±0.418 stderr; $0.01457/test |
| LiveCodeBench | 88.2% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.941 stderr; $0.040/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -2.8 | HighGrok 4.6 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 51/56; Thurstone comparison score (centered at 0); est. win chance 13%; 95% bootstrap -2.977 to -2.700 |
| LMArena Code Arena (WebDev) | 1620 | Highgrok-4.6-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 20 (CI rank 15-24); 95% CI 1612-1628; 8120 votes |
| SciCode | 49.4% | LowGrok 4.6 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated) |
| SciCode | 55.9% | MediumGrok 4.6 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated) |
| SciCode | 56.5% | HighGrok 4.6 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated) |
| SciCode | 53.0% | Extra highGrok 4.6 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated) |
| SimpleBench | 75.9% | DefaultGrok 4.6 | Independent testIndependentSimpleBench ↗ | 13 Aug 2026 | — | AVG@5, temp 0.7; rank 13th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 49.3% | Highgrok-4.6_high | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SimpleQA Verified | 48.9% | Extra highgrok-4.6_xhigh | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SkillsBench | 55.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.844 stderr; $0.921/test |
| SWE Atlas - Codebase QnA | 58.3% | Extra highGrok 4.6 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Grok Build | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE-bench Verified | 95.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±0.918 stderr; $0.785/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 75.3% | LowGrok 4.6 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 84.3% | MediumGrok 4.6 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 88.4% | HighGrok 4.6 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 78.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±2.085 stderr; $0.454/test |
| Terminal-Bench 2.1 | 88.0% | Extra highGrok 4.6 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated) |
| Terminal-Bench 3.0 | 26.5% | High | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Grok Build | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 2.9B, run cost $2.1k |
| Terminal-Bench 3.0 | 26.0% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | |
| Terminal-Bench 4.0 | 3.0% | LowGrok 4.6 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-low) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 13.1% | MediumGrok 4.6 (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-medium) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 21.2% | HighGrok 4.6 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 20.3% | High | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Grok Build | TB 4.0.0 (66 tasks); ±3.09 95% CI; 330 trials; total run cost $3592; model release 2026-08-12 |
| Terminal-Bench 4.0 | 17.2% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.336 stderr; $5.071/test |
| Terminal-Bench 4.0 | 20.3% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | Comparison column in the Grok 4.7 launch post. |
| Terminal-Bench 4.0 | 17.7% | Extra highGrok 4.6 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Grok Build | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 17.2% | Extra highGrok 4.6 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-6-xhigh) (AA marks this variant deprecated) |
| Vals Code Migration | 44.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.479 stderr; $14.575/test |
| Vals Index | 52.1% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.153 stderr; $4.493/test |
| Vals Legal Research Bench | 48.1% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id grok/grok-4.6; rank 8/72; ±3.473 stderr; $1.530204/test |
| Vals Public Benefits Bench v1.1 | 66.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id grok/grok-4.6; rank 15/45; ±1.225 stderr; $0.91559/test |
| Vending-Bench 2 | $9,047 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$1,604; rank 9; only top 10 rendered server-side (57 more behind 'Show more') |