Skip to content
Bencher

Models · Z.ai (Zhipu) · Out sinceReleased 7 Apr 2026

GLM-5.1

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

GLM-5.1 is made by Z.ai (Zhipu). We don't have enough test results yet to rank it. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.

Previous flagship (thinking on/off). Context/max output from OpenRouter.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Mid-priced$1.40 / $4.40
Price per 1M (blended)Blended / 1M
$2.15
MemoryContext
205K
Longest answerMax output
131K
Test resultsResults
18 (4 independent4 indep.)
Out sinceReleased
7 Apr 2026
Made byVendor
Z.ai (Zhipu)
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as z-ai/glm-5.1.

Route it as z-ai/glm-5.1 at $0.96 in / $3.03 out per 1M tokens, 205K context. Listed since 7 Apr 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

Details comingselfHost data pending

This model can be downloaded and run on your own servers. We're collecting its licence, size and hardware needs now.Open-weights model; licence, parameter count, hardware tier and base-model availability will appear after the next data run.

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AIME 202695.3%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
BrowseComp68.0%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—79.3 with context management
CyberGym68.7%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
GPQA Diamond86.2%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
HMMT February 202682.6%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—HMMT Feb. 2026
Humanity's Last Exam31.0%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
Humanity's Last Exam (with tools)52.3%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
IMO-AnswerBench83.8%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
LLM Creative Story-Writing Benchmark (Lech Mazur)-1DefaultGLM-5.1Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 39/56; Thurstone comparison score (centered at 0); est. win chance 35%; 95% bootstrap -1.229 to -0.844
MCP Atlas75.6%Defaultglm-5p1Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±2.7; entry added 2025-12-17
MCP Atlas71.8%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—Public set
NL2Repo-Bench42.7%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
SWE-Bench Pro (public, v1)58.4%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
SWE-bench Verified74.2%Defaultglm-5.1Independent testIndependentEpoch AI ↗15 May 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 2.0pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)
tau3-bench70.6%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—
Terminal-Bench 2.063.5%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026Terminus 269.0 with Claude Code (best self-reported)
Terminal-Bench 2.158.7%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 2.1 (89 tasks, archived); effort not listed
Toolathlon40.7%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗7 Apr 2026—Reported as 'Tool-Decathlon'