Models · SpaceXAI (formerly xAI) · Out sinceReleased 8 Jul 2026
Grok 4.5#22 for coding.#22 for coding, best at high effort.
Grok 4.5 is made by SpaceXAI (formerly xAI). Among the models we track it ranks #22 for coding. It's mid-priced to use.
API release 2026-07-08 (release notes); news post 2026-07-16. Default effort high (launch notes listed low/medium/high; model page now lists xhigh too). Alias grok-build-latest now points to grok-4.5. Trained alongside Cursor.
- —
- —
- 41.4 / 100 · #22
- Mid-priced$2 / $6
- $3
- 500K
- —
- 39 (24 independent24 indep.)
- 8 Jul 2026
- SpaceXAI (formerly xAI)
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (High thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as x-ai/grok-4.5.
Route it as x-ai/grok-4.5 at $2 in / $6 out per 1M tokens, 500K context. Listed since 8 Jul 2026.
include_reasoninglogprobsmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstemperaturetool_choicetoolstop_logprobstop_p
Thinking level
Reasoning effort
How long should Grok 4.5Where Grok 4.5think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 3.0
Best at HighBest at High
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase | 1313 | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Comparison column in the Grok 4.6 launch post. |
| APEX-Agents | 47.1% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Comparison column in the Grok 4.6 launch post. |
| APEX-SWE | 53.6% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Comparison column in the Grok 4.6 launch post. |
| Artificial Analysis Intelligence Index | 38.8 | HighGrok 4.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug grok-4-5; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $1.04/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 56 | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Comparison column in the Grok 4.6 launch post. |
| Artificial Analysis output speed | 55 tok/s | HighGrok 4.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 12.6s; list price $2/6 per 1M in/out (AA marks this variant deprecated) |
| CursorBench 3.2.0 | 66.7% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Comparison column in the Grok 4.6 launch post. |
| DeepSWE | 54.0% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Integer as published. Comparison column in the Grok 4.6 launch post. |
| DeepSWE | 53.0% | Defaulteffort not stated (API default high) | Maker's own figureVendor-reportedSpaceXAI ↗ | 16 Jul 2026 | mini-swe-agent (run by Datacurve) | Integer as published. |
| DeepSWE 1.0 | 62.0% | Defaulteffort not stated (API default high) | Maker's own figureVendor-reportedSpaceXAI ↗ | 16 Jul 2026 | each provider's harness, run by Artificial Analysis | Eval created by Datacurve. |
| FrontierCode | 42.4% | Defaultgrok-4.5_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | grok-build | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness grok-build; Mean@5 |
| FrontierCode v1.1 (Extended) | 56.6% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Comparison column in the Grok 4.6 launch post. |
| GDPval-AA v2 | 1526 | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Comparison column in the Grok 4.6 launch post. |
| GPQA Diamond | 93.4% | Highgrok-4.5_high | Independent testIndependentEpoch AI ↗ | 8 Jul 2026 | — | Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 93.1% | HighGrok 4.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated) |
| GPQA Diamond | 92.9% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.288 stderr; $0.044/test |
| Harvey's Legal Agent Benchmark (Vals) | 12.9% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Harvey LAB (Vals). Comparison column in the Grok 4.6 launch post. |
| Humanity's Last Exam | 42.7% | HighGrok 4.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated) |
| LegalBench (Vals) | 86.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id grok/grok-4.5; rank 17/149; ±0.413 stderr; $0.004535/test |
| LiveCodeBench | 87.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.961 stderr; $0.045/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -5.1 | HighGrok 4.5 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 56/56; Thurstone comparison score (centered at 0); est. win chance 2%; 95% bootstrap -5.160 to -4.974 |
| LMArena Search Arena | 1213 | Defaultgrok-4.5 | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 8 (rank range 6-13); 95% CI 1205.8-1219.5; 31505 votes |
| SciCode | 55.0% | HighGrok 4.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated) |
| SimpleBench | 70.0% | DefaultGrok 4.5 | Independent testIndependentSimpleBench ↗ | 8 Jul 2026 | — | AVG@5, temp 0.7; rank 18th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 48.3% | Highgrok-4.5_high | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SkillsBench | 66.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.596 stderr; $0.578/test |
| SWE Marathon | 29.0% | Defaulteffort not stated (API default high) | Maker's own figureVendor-reportedSpaceXAI ↗ | 16 Jul 2026 | — | Resolution rate, pass@1. |
| SWE-Bench Pro (public, v1) | 64.7% | Defaulteffort not stated (API default high) | Maker's own figureVendor-reportedSpaceXAI ↗ | 16 Jul 2026 | — | Resolve rate; avg 15,954 output tokens per task (4.2x fewer than Opus 4.8 max). |
| SWE-bench Verified | 86.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.525 stderr; $0.540/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| SWE-rebench | 63.8% | HighGrok 4.5 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | 1 Oct 2026 | SWE-rebench standard scaffold | time window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±0.60; pass@5 77.5%; $1.47/problem |
| Terminal-Bench 2.1 | 81.6% | HighGrok 4.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 79.3% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Cursor CLI | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 83.3% | Defaulteffort not stated (API default high) | Maker's own figureVendor-reportedSpaceXAI ↗ | 16 Jul 2026 | — | |
| Terminal-Bench 3.0 | 15.7% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | 12 Aug 2026 | — | Comparison column in the Grok 4.6 launch post. |
| Terminal-Bench 3.0 | 15.7% | Extra high | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Cursor CLI | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 1.2B, run cost $766 |
| Terminal-Bench 4.0 | 12.4% | High | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Grok Build | TB 4.0.0 (66 tasks); ±2.62 95% CI; 330 trials; total run cost $2094; model release 2026-07-16 |
| Terminal-Bench 4.0 | 10.6% | HighGrok 4.5 (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug grok-4-5) (AA marks this variant deprecated) |
| Vals CorpFin v2 | 67.4% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id grok/grok-4.5; rank 13/134; ±0.924 stderr; $0.221291/test |
| Vals SRE Bench | 0.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±0.539 stderr; $13.415/test |