Skip to content
Bencher

Models · SpaceXAI (formerly xAI) · Out sinceReleased 8 Jul 2026

Grok 4.5#22 for coding.#22 for coding, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Grok 4.5 is made by SpaceXAI (formerly xAI). Among the models we track it ranks #22 for coding. It's mid-priced to use.

API release 2026-07-08 (release notes); news post 2026-07-16. Default effort high (launch notes listed low/medium/high; model page now lists xhigh too). Alias grok-build-latest now points to grok-4.5. Trained alongside Cursor.

Writing & creativity
—
Research & analysis
—
Coding
41.4 / 100 · #22
Price
Mid-priced$2 / $6
Price per 1M (blended)Blended / 1M
$3
MemoryContext
500K
Longest answerMax output
—
Test resultsResults
39 (24 independent24 indep.)
Out sinceReleased
8 Jul 2026
Made byVendor
SpaceXAI (formerly xAI)
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra high

Highlighted: where it did best for coding (High thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as x-ai/grok-4.5.

Route it as x-ai/grok-4.5 at $2 in / $6 out per 1M tokens, 500K context. Listed since 8 Jul 2026.

include_reasoninglogprobsmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstemperaturetool_choicetoolstop_logprobstop_p

Thinking level

Reasoning effort

How long should Grok 4.5Where Grok 4.5think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 3.0

Best at HighBest at High

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase1313HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Comparison column in the Grok 4.6 launch post.
APEX-Agents47.1%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Comparison column in the Grok 4.6 launch post.
APEX-SWE53.6%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Comparison column in the Grok 4.6 launch post.
Artificial Analysis Intelligence Index38.8HighGrok 4.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug grok-4-5; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $1.04/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index56HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Comparison column in the Grok 4.6 launch post.
Artificial Analysis output speed55 tok/sHighGrok 4.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 12.6s; list price $2/6 per 1M in/out (AA marks this variant deprecated)
CursorBench 3.2.066.7%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Comparison column in the Grok 4.6 launch post.
DeepSWE54.0%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Integer as published. Comparison column in the Grok 4.6 launch post.
DeepSWE53.0%Defaulteffort not stated (API default high)Maker's own figureVendor-reportedSpaceXAI ↗16 Jul 2026mini-swe-agent (run by Datacurve)Integer as published.
DeepSWE 1.062.0%Defaulteffort not stated (API default high)Maker's own figureVendor-reportedSpaceXAI ↗16 Jul 2026each provider's harness, run by Artificial AnalysisEval created by Datacurve.
FrontierCode42.4%Defaultgrok-4.5_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026grok-buildread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness grok-build; Mean@5
FrontierCode v1.1 (Extended)56.6%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Comparison column in the Grok 4.6 launch post.
GDPval-AA v21526HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Comparison column in the Grok 4.6 launch post.
GPQA Diamond93.4%Highgrok-4.5_highIndependent testIndependentEpoch AI ↗8 Jul 2026—Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond93.1%HighGrok 4.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated)
GPQA Diamond92.9%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±1.288 stderr; $0.044/test
Harvey's Legal Agent Benchmark (Vals)12.9%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Harvey LAB (Vals). Comparison column in the Grok 4.6 launch post.
Humanity's Last Exam42.7%HighGrok 4.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated)
LegalBench (Vals)86.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id grok/grok-4.5; rank 17/149; ±0.413 stderr; $0.004535/test
LiveCodeBench87.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±0.961 stderr; $0.045/test
LLM Creative Story-Writing Benchmark (Lech Mazur)-5.1HighGrok 4.5 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 56/56; Thurstone comparison score (centered at 0); est. win chance 2%; 95% bootstrap -5.160 to -4.974
LMArena Search Arena1213Defaultgrok-4.5Independent testIndependentLMArena ↗24 Aug 2026—rank 8 (rank range 6-13); 95% CI 1205.8-1219.5; 31505 votes
SciCode55.0%HighGrok 4.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated)
SimpleBench70.0%DefaultGrok 4.5Independent testIndependentSimpleBench ↗8 Jul 2026—AVG@5, temp 0.7; rank 18th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified48.3%Highgrok-4.5_highIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SkillsBench66.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.596 stderr; $0.578/test
SWE Marathon29.0%Defaulteffort not stated (API default high)Maker's own figureVendor-reportedSpaceXAI ↗16 Jul 2026—Resolution rate, pass@1.
SWE-Bench Pro (public, v1)64.7%Defaulteffort not stated (API default high)Maker's own figureVendor-reportedSpaceXAI ↗16 Jul 2026—Resolve rate; avg 15,954 output tokens per task (4.2x fewer than Opus 4.8 max).
SWE-bench Verified86.6%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.525 stderr; $0.540/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
SWE-rebench63.8%HighGrok 4.5 [high]Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±0.60; pass@5 77.5%; $1.47/problem
Terminal-Bench 2.181.6%HighGrok 4.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated)
Terminal-Bench 2.179.3%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Cursor CLITB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.183.3%Defaulteffort not stated (API default high)Maker's own figureVendor-reportedSpaceXAI ↗16 Jul 2026—
Terminal-Bench 3.015.7%HighMaker's own figureVendor-reportedSpaceXAI ↗12 Aug 2026—Comparison column in the Grok 4.6 launch post.
Terminal-Bench 3.015.7%Extra highIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Cursor CLITB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 1.2B, run cost $766
Terminal-Bench 4.012.4%HighIndependent testIndependentTerminal-Bench ↗1 Oct 2026Grok BuildTB 4.0.0 (66 tasks); ±2.62 95% CI; 330 trials; total run cost $2094; model release 2026-07-16
Terminal-Bench 4.010.6%HighGrok 4.5 (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated)
Vals CorpFin v267.4%Highreasoning_effort=highIndependent testIndependentVals.ai ↗12 Aug 2026—vals id grok/grok-4.5; rank 13/134; ±0.924 stderr; $0.221291/test
Vals SRE Bench0.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±0.539 stderr; $13.415/test