Models · SpaceXAI (formerly xAI) · Out sinceReleased 21 Sep 2026
Grok 4.7#15 for coding.#15 for coding, best at extra high effort.
Grok 4.7 is made by SpaceXAI (formerly xAI). Among the models we track it ranks #15 for coding, #25 for research and analysis, #32 for writing. It's mid-priced to use.
SpaceXAI flagship (xAI was acquired by SpaceX in 2026). Default reasoning effort = high. >200k-token prompts $4/$12. 'No text output limit' per release notes. Grok 4.7 Fast (same model, 2x price/2x speed) only in Cursor and Grok Build. Knowledge cutoff May 2026. No model card found.
- 29.2 / 100 · #32
- 60.5 / 100 · #25
- 53.0 / 100 · #15
- Mid-priced$2 / $6
- $3
- 500K
- —
- 65 (57 independent57 indep.)
- 21 Sep 2026
- SpaceXAI (formerly xAI)
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Extra high thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as x-ai/grok-4.7.
Route it as x-ai/grok-4.7 at $2 in / $6 out per 1M tokens, 500K context. Listed since 21 Sep 2026.
include_reasoninglogprobsmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstemperaturetool_choicetoolstop_logprobstop_p
Thinking level
Reasoning effort
How long should Grok 4.7Where Grok 4.7think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at Extra highBest at Extra high
CursorBench
Best at Extra highBest at Extra high
DeepSWE
Best at Extra highBest at Extra high
SciCode
Best at High, worse abovePeaks at High
Artificial Analysis Intelligence Index
Best at Extra highBest at Extra high
AA-LCR
Best at High, worse abovePeaks at High
AA-Omniscience Accuracy
Best at High, worse abovePeaks at High
AA-Omniscience Hallucination Rate
Best at Extra highBest at Extra high
Lower is better on this test
AA-Omniscience Index
Best at Extra highBest at Extra high
GDPval-AA v2.1
Best at Extra highBest at Extra high
Humanity's Last Exam
Best at Extra highBest at Extra high
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1657 | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1657 | Extra high | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | |
| AA-LCR | 77.0% | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 76.7% | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 47.8% | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 47.5% | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 32.4% | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 29.3% | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 30.9 | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 32 | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Artificial Analysis Coding Agent Index | 56.3 | Extra highGrok 4.7 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Grok Build | agent Grok Build; components: DeepSWE v1.1 72.6, SWE-Atlas-QnA 62.9, Terminal-Bench v4 33.3; avg cost $8.82/task; avg wall time 39 min/task |
| Artificial Analysis Intelligence Index | 46.3 | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug grok-4-7-high; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $2.73/task |
| Artificial Analysis Intelligence Index | 46.4 | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug grok-4-7; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $3.74/task |
| Artificial Analysis output speed | 73 tok/s | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 31.7s; list price $2/6 per 1M in/out |
| Artificial Analysis output speed | 73 tok/s | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 54.7s; list price $2/6 per 1M in/out |
| CursorBench | 33.1% | LowGrok 4.7 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 44; $1.58/task; 15,677 tokens/task; 40 steps/task |
| CursorBench | 41.6% | MediumGrok 4.7 Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 22; $3.49/task; 36,683 tokens/task; 60 steps/task |
| CursorBench | 43.9% | HighGrok 4.7 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 17; $4.69/task; 56,382 tokens/task; 71 steps/task |
| CursorBench | 46.3% | Extra highGrok 4.7 Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 13; $6.01/task; 70,141 tokens/task; 88 steps/task |
| CursorBench | 46.3% | Extra high | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | |
| DeepSWE | 71.0% | Highhigh effort | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | Post table header is 'Grok 4.7 xHigh' but this DeepSWE score is footnoted '* high effort'. |
| DeepSWE | 72.6% | Extra highGrok 4.7 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Grok Build | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| EEBench | 64.0% | Extra high | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | |
| EQ-Bench Creative Writing v3 (Elo) | 2007 | Default*grok-4.7 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 8; rubric score 17.16/20; slop 8.47; avg length 6282 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*) |
| FrontierSWE | 29.5% | Default | Independent testIndependentFrontierSWE ↗ | 1 Oct 2026 | proximus | FrontierSWE V2, mean@5 over 34 tasks (20h budget); ±12.6; $318.77/trial; 12.1h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort |
| GDPval-AA (version unstated) | 1695 | Extra high | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | Elo; version not stated in post. Reported by xAI as "GDPval"; Elo scale implies GDPval-AA. |
| GDPval-AA v2.1 | 1694 | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1678.25-1709.19 |
| GDPval-AA v2.1 | 1695 | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1675.39-1715.03 |
| Harvey's Legal Agent Benchmark (Vals) | 19.6% | Extra high | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | |
| HealthBench Professional | 56.7% | Extra high | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | |
| HLE Diamond | 23.4% | Defaultgrok-4.7 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 9 (Scale rank accounts for CI); ±2.6; entry added 2026-02-17 |
| Humanity's Last Exam | 42.3% | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-7-high) |
| Humanity's Last Exam | 43.1% | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-7) |
| IOI (Vals) | 57.7% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±1.987 stderr; $12.708/test |
| LegalBench (Vals) | 84.4% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id grok/grok-4.7; rank 32/149; ±0.456 stderr; $0.010658/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.4 | HighGrok 4.7 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 23/56; Thurstone comparison score (centered at 0); est. win chance 57%; 95% bootstrap 0.246 to 0.655 |
| LMArena Code Arena (WebDev) | 1636 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 15 (CI rank 13-23); 95% CI 1624-1648; 3062 votes |
| LMArena Text - Creative Writing | 1433 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 68 (rank range 27-99); 95% CI 1415.6-1450.2; 1293 votes; style-controlled |
| LMArena Text - Expert | 1459 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 101 (rank range 47-158); 95% CI 1434.7-1482.5; 580 votes; style-controlled |
| LMArena Text - Hard Prompts | 1462 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 88 (rank range 66-118); 95% CI 1451.7-1472.2; 3405 votes; style-controlled |
| LMArena Text - Instruction Following | 1442 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 77 (rank range 45-109); 95% CI 1428.3-1455.5; 1875 votes; style-controlled |
| LMArena Text - Longer Query | 1451 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 87 (rank range 57-124); 95% CI 1438.0-1463.1; 2342 votes; style-controlled |
| LMArena Text - Multi-Turn | 1425 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 122 (rank range 79-170); 95% CI 1404.0-1445.4; 816 votes; style-controlled |
| LMArena Text - Non-English | 1435 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 82 (rank range 54-105); 95% CI 1425.1-1445.5; 3428 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1430 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 110 (rank range 62-156); 95% CI 1411.5-1448.6; 996 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1460 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 74 (rank range 17-143); 95% CI 1433.6-1485.7; 510 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1427 | Extra highgrok-4.7-xhigh | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 82 (rank range 53-118); 95% CI 1411.4-1442.2; 1579 votes; style-controlled |
| PRBench Legal (Scale) | 47.6% | Extra highGrok 4.7 (xHigh) | Independent testIndependentScale AI (SEAL) ↗ | 25 Sep 2026 | — | Scale rank 9; ±1.79 CI |
| ProgramBench (fully resolved) | 0.5% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.5 stderr; $46.491/test; strict fully-resolved rate |
| SciCode | 57.8% | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-7-high) |
| SciCode | 57.4% | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-7) |
| SWE Atlas - Codebase QnA | 62.9% | Extra highGrok 4.7 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Grok Build | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| Terminal-Bench 2.1 | 73.4% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.498 stderr; $1.077/test |
| Terminal-Bench 4.0 | 24.7% | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-7-high) |
| Terminal-Bench 4.0 | 37.6% | Extra high | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Grok Build | TB 4.0.0 (66 tasks); ±3.54 95% CI; 330 trials; total run cost $3683; model release 2026-09-21 |
| Terminal-Bench 4.0 | 33.3% | Extra highGrok 4.7 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Grok Build | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 28.8% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±2.314 stderr; $18.092/test |
| Terminal-Bench 4.0 | 25.8% | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-7) |
| Terminal-Bench 4.0 | 37.6% | Extra high | Maker's own figureVendor-reportedSpaceXAI ↗ | 21 Sep 2026 | — | |
| Vals Code Migration | 44.8% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.215 stderr; $36.552/test |
| Vals Index | 55.0% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.072 stderr; $12.118/test |
| Vals Legal Research Bench | 47.1% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id grok/grok-4.7; rank 12/72; ±3.469 stderr; $4.920812/test |
| Vals Public Benefits Bench v1.1 | 65.6% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id grok/grok-4.7; rank 18/45; ±1.235 stderr; $2.246775/test |
| Vals Vibe Code Bench | 86.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±2.181 stderr; $15.827/test |
| Vending-Bench 2 | $10,537 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$652; rank 6; only top 10 rendered server-side (57 more behind 'Show more') |