Models · Z.ai (Zhipu) · Out sinceReleased 7 Apr 2026
GLM-5.1
GLM-5.1 is made by Z.ai (Zhipu). We don't have enough test results yet to rank it. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.
Previous flagship (thinking on/off). Context/max output from OpenRouter.
- —
- —
- —
- Mid-priced$1.40 / $4.40
- $2.15
- 205K
- 131K
- 18 (4 independent4 indep.)
- 7 Apr 2026
- Z.ai (Zhipu)
- text
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as z-ai/glm-5.1.
Route it as z-ai/glm-5.1 at $0.96 in / $3.03 out per 1M tokens, 205K context. Listed since 7 Apr 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
Details comingselfHost data pending
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AIME 2026 | 95.3% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| BrowseComp | 68.0% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | 79.3 with context management |
| CyberGym | 68.7% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| GPQA Diamond | 86.2% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| HMMT February 2026 | 82.6% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | HMMT Feb. 2026 |
| Humanity's Last Exam | 31.0% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| Humanity's Last Exam (with tools) | 52.3% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| IMO-AnswerBench | 83.8% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -1 | DefaultGLM-5.1 | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 39/56; Thurstone comparison score (centered at 0); est. win chance 35%; 95% bootstrap -1.229 to -0.844 |
| MCP Atlas | 75.6% | Defaultglm-5p1 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±2.7; entry added 2025-12-17 |
| MCP Atlas | 71.8% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | Public set |
| NL2Repo-Bench | 42.7% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| SWE-Bench Pro (public, v1) | 58.4% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| SWE-bench Verified | 74.2% | Defaultglm-5.1 | Independent testIndependentEpoch AI ↗ | 15 May 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 2.0pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| tau3-bench | 70.6% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | |
| Terminal-Bench 2.0 | 63.5% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | Terminus 2 | 69.0 with Claude Code (best self-reported) |
| Terminal-Bench 2.1 | 58.7% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 2.1 (89 tasks, archived); effort not listed |
| Toolathlon | 40.7% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | 7 Apr 2026 | — | Reported as 'Tool-Decathlon' |