Skip to content
Bencher

Models · SpaceXAI (formerly xAI) · Out sinceReleased 21 Sep 2026

Grok 4.7#15 for coding.#15 for coding, best at extra high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Grok 4.7 is made by SpaceXAI (formerly xAI). Among the models we track it ranks #15 for coding, #25 for research and analysis, #32 for writing. It's mid-priced to use.

SpaceXAI flagship (xAI was acquired by SpaceX in 2026). Default reasoning effort = high. >200k-token prompts $4/$12. 'No text output limit' per release notes. Grok 4.7 Fast (same model, 2x price/2x speed) only in Cursor and Grok Build. Knowledge cutoff May 2026. No model card found.

Writing & creativity
29.2 / 100 · #32
Research & analysis
60.5 / 100 · #25
Coding
53.0 / 100 · #15
Price
Mid-priced$2 / $6
Price per 1M (blended)Blended / 1M
$3
MemoryContext
500K
Longest answerMax output
—
Test resultsResults
65 (57 independent57 indep.)
Out sinceReleased
21 Sep 2026
Made byVendor
SpaceXAI (formerly xAI)
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra high

Highlighted: where it did best for coding (Extra high thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as x-ai/grok-4.7.

Route it as x-ai/grok-4.7 at $2 in / $6 out per 1M tokens, 500K context. Listed since 21 Sep 2026.

include_reasoninglogprobsmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstemperaturetool_choicetoolstop_logprobstop_p

Thinking level

Reasoning effort

How long should Grok 4.7Where Grok 4.7think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at Extra highBest at Extra high

CursorBench

Best at Extra highBest at Extra high

DeepSWE

Best at Extra highBest at Extra high

SciCode

Best at High, worse abovePeaks at High

Artificial Analysis Intelligence Index

Best at Extra highBest at Extra high

AA-LCR

Best at High, worse abovePeaks at High

AA-Omniscience Accuracy

Best at High, worse abovePeaks at High

AA-Omniscience Hallucination Rate

Best at Extra highBest at Extra high

Lower is better on this test

AA-Omniscience Index

Best at Extra highBest at Extra high

GDPval-AA v2.1

Best at Extra highBest at Extra high

Humanity's Last Exam

Best at Extra highBest at Extra high

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11657Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11657Extra highMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—
AA-LCR77.0%HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR76.7%Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy47.8%HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy47.5%Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate32.4%HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate29.3%Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index30.9HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index32Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Artificial Analysis Coding Agent Index56.3Extra highGrok 4.7 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Grok Buildagent Grok Build; components: DeepSWE v1.1 72.6, SWE-Atlas-QnA 62.9, Terminal-Bench v4 33.3; avg cost $8.82/task; avg wall time 39 min/task
Artificial Analysis Intelligence Index46.3HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug grok-4-7-high; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $2.73/task
Artificial Analysis Intelligence Index46.4Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug grok-4-7; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $3.74/task
Artificial Analysis output speed73 tok/sHighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 31.7s; list price $2/6 per 1M in/out
Artificial Analysis output speed73 tok/sExtra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 54.7s; list price $2/6 per 1M in/out
CursorBench33.1%LowGrok 4.7 LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 44; $1.58/task; 15,677 tokens/task; 40 steps/task
CursorBench41.6%MediumGrok 4.7 MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 22; $3.49/task; 36,683 tokens/task; 60 steps/task
CursorBench43.9%HighGrok 4.7 HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 17; $4.69/task; 56,382 tokens/task; 71 steps/task
CursorBench46.3%Extra highGrok 4.7 Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 13; $6.01/task; 70,141 tokens/task; 88 steps/task
CursorBench46.3%Extra highMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—
DeepSWE71.0%Highhigh effortMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—Post table header is 'Grok 4.7 xHigh' but this DeepSWE score is footnoted '* high effort'.
DeepSWE72.6%Extra highGrok 4.7 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Grok BuildAA Coding Agent Index component; pass@1 avg of 3 attempts
EEBench64.0%Extra highMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—
EQ-Bench Creative Writing v3 (Elo)2007Default*grok-4.7Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 8; rubric score 17.16/20; slop 8.47; avg length 6282 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*)
FrontierSWE29.5%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±12.6; $318.77/trial; 12.1h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort
GDPval-AA (version unstated)1695Extra highMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—Elo; version not stated in post. Reported by xAI as "GDPval"; Elo scale implies GDPval-AA.
GDPval-AA v2.11694HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1678.25-1709.19
GDPval-AA v2.11695Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1675.39-1715.03
Harvey's Legal Agent Benchmark (Vals)19.6%Extra highMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—
HealthBench Professional56.7%Extra highMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—
HLE Diamond23.4%Defaultgrok-4.7Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 9 (Scale rank accounts for CI); ±2.6; entry added 2026-02-17
Humanity's Last Exam42.3%HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-7-high)
Humanity's Last Exam43.1%Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-7)
IOI (Vals)57.7%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—±1.987 stderr; $12.708/test
LegalBench (Vals)84.4%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—vals id grok/grok-4.7; rank 32/149; ±0.456 stderr; $0.010658/test
LLM Creative Story-Writing Benchmark (Lech Mazur)0.4HighGrok 4.7 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 23/56; Thurstone comparison score (centered at 0); est. win chance 57%; 95% bootstrap 0.246 to 0.655
LMArena Code Arena (WebDev)1636Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 15 (CI rank 13-23); 95% CI 1624-1648; 3062 votes
LMArena Text - Creative Writing1433Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 68 (rank range 27-99); 95% CI 1415.6-1450.2; 1293 votes; style-controlled
LMArena Text - Expert1459Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 101 (rank range 47-158); 95% CI 1434.7-1482.5; 580 votes; style-controlled
LMArena Text - Hard Prompts1462Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 88 (rank range 66-118); 95% CI 1451.7-1472.2; 3405 votes; style-controlled
LMArena Text - Instruction Following1442Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 77 (rank range 45-109); 95% CI 1428.3-1455.5; 1875 votes; style-controlled
LMArena Text - Longer Query1451Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 87 (rank range 57-124); 95% CI 1438.0-1463.1; 2342 votes; style-controlled
LMArena Text - Multi-Turn1425Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 122 (rank range 79-170); 95% CI 1404.0-1445.4; 816 votes; style-controlled
LMArena Text - Non-English1435Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 82 (rank range 54-105); 95% CI 1425.1-1445.5; 3428 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1430Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 110 (rank range 62-156); 95% CI 1411.5-1448.6; 996 votes; style-controlled
LMArena Text - Occupational: Legal & Government1460Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 74 (rank range 17-143); 95% CI 1433.6-1485.7; 510 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1427Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗30 Sep 2026—rank 82 (rank range 53-118); 95% CI 1411.4-1442.2; 1579 votes; style-controlled
PRBench Legal (Scale)47.6%Extra highGrok 4.7 (xHigh)Independent testIndependentScale AI (SEAL) ↗25 Sep 2026—Scale rank 9; ±1.79 CI
ProgramBench (fully resolved)0.5%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.5 stderr; $46.491/test; strict fully-resolved rate
SciCode57.8%HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-7-high)
SciCode57.4%Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-7)
SWE Atlas - Codebase QnA62.9%Extra highGrok 4.7 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Grok BuildAA Coding Agent Index component; pass@1 avg of 3 attempts
Terminal-Bench 2.173.4%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.498 stderr; $1.077/test
Terminal-Bench 4.024.7%HighGrok 4.7 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-7-high)
Terminal-Bench 4.037.6%Extra highIndependent testIndependentTerminal-Bench ↗1 Oct 2026Grok BuildTB 4.0.0 (66 tasks); ±3.54 95% CI; 330 trials; total run cost $3683; model release 2026-09-21
Terminal-Bench 4.033.3%Extra highGrok 4.7 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Grok BuildAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.028.8%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±2.314 stderr; $18.092/test
Terminal-Bench 4.025.8%Extra highGrok 4.7 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-7)
Terminal-Bench 4.037.6%Extra highMaker's own figureVendor-reportedSpaceXAI ↗21 Sep 2026—
Vals Code Migration44.8%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—±4.215 stderr; $36.552/test
Vals Index55.0%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗30 Sep 2026—±1.072 stderr; $12.118/test
Vals Legal Research Bench47.1%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—vals id grok/grok-4.7; rank 12/72; ±3.469 stderr; $4.920812/test
Vals Public Benefits Bench v1.165.6%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—vals id grok/grok-4.7; rank 18/45; ±1.235 stderr; $2.246775/test
Vals Vibe Code Bench86.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±2.181 stderr; $15.827/test
Vending-Bench 2$10,537DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$652; rank 6; only top 10 rendered server-side (57 more behind 'Show more')