Models · Z.ai (Zhipu) · Out sinceReleased 16 Jun 2026
GLM-5.2#20 for coding.#20 for coding, best at its default setting.
GLM-5.2 is made by Z.ai (Zhipu). Among the models we track it ranks #20 for coding. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.
MIT license, 1M context, IndexShare sparse attention. Introduced effort levels (High/Max exposed in Coding Plan).
- —
- —
- 49.8 / 100 · #20
- Mid-priced$1.40 / $4.40
- $2.15
- 1M
- 131K
- 51 (26 independent26 indep.)
- 16 Jun 2026
- Z.ai (Zhipu)
- text
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (its standard thinking level).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as z-ai/glm-5.2.
Route it as z-ai/glm-5.2 at $1.40 in / $4.40 out per 1M tokens, 1M context. Listed since 16 Jun 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
Details comingselfHost data pending
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Agents' Last Exam | 23.8% | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | GLM-5.2 column in the GLM-5.3 launch table; ALE-CLI |
| AIME 2026 | 99.2% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | |
| Artificial Analysis Intelligence Index | 33.7 | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug glm-5-2; list price $1.4/4.4 per 1M in/out; cost to run AA Intelligence Index $1.47/task (AA marks this variant deprecated) |
| Artificial Analysis output speed | 82 tok/s | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 3.1s; list price $1.4/4.4 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 26.2% | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | GLM-5.2 column in the GLM-5.3 launch table; v1.0.6 |
| CyberGym | 77.2% | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | GLM-5.2 column in the GLM-5.3 launch table |
| DeepSWE | 46.2% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | mini-swe-agent | official pier framework, 2h timeout |
| Design Arena (all categories) | 1297 | Defaultglm-5.2 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±1.9 SE; 42958 battles; win rate 52.1% |
| Design Arena (fullstack) | 1239 | Defaultglm-5.2 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±7.3 SE; 2758 battles; win rate 60.5% |
| EQ-Bench Creative Writing v3 (Elo) | 1757 | Defaultzai-org/GLM-5.2 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 29; rubric score 16.44/20; slop 13.11; avg length 6104 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| ExploitBench | 24.4% | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | GLM-5.2 column in the GLM-5.3 launch table |
| FrontierSWE | 74.4% | Maxmax effort | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | Run by Proximal at max effort; dominance as of 2026-06-16 |
| GDPval-AA v2 | 1508 | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | GLM-5.2 column in the GLM-5.3 launch table |
| GPQA Diamond | 89.5% | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated) |
| GPQA Diamond | 91.2% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | |
| HMMT February 2026 | 92.5% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | HMMT Feb. 2026 |
| Humanity's Last Exam | 41.1% | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated) |
| Humanity's Last Exam | 40.5% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | Text-only subset |
| Humanity's Last Exam (with tools) | 54.7% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | 300K ctx, no context management |
| IMO-AnswerBench | 91.0% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.5 | MaxGLM-5.2 (max) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 22/56; Thurstone comparison score (centered at 0); est. win chance 57%; 95% bootstrap 0.380 to 0.561 |
| LMArena Code Arena (WebDev) | 1605 | Maxglm-5.2-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 25 (CI rank 22-26); 95% CI 1598-1611; 14543 votes |
| MCP Atlas | 77.8% | Defaultglm-5p2 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±2.6; entry added 2026-07-20 |
| MCP Atlas | 76.8% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | Public set, think mode |
| NL2Repo-Bench | 48.9% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | 400K ctx |
| PostTrainBench v1.1 | 34.3% | Maxmax effort | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | Run by PostTrainBench at max effort |
| ProgramBench (Almost Solved) | 9.5% | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | GLM-5.2 column in the GLM-5.3 launch table; ProgramBench 'Almost Solved' |
| ProgramBench (avg test pass rate) | 64.6% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 8.5%; avg cost $25.36/task |
| ProgramBench (fully resolved) | 0.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.5 stderr; $12.577/test; strict fully-resolved rate |
| ProgramBench (fully resolved) | 63.7% | Maxmax effort | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | Claude Code 2.1.156 | 200 instances, reasoning_effort=max |
| ProgramBench (fully resolved) | 0.0% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 8.5%; avg cost $25.36/task |
| SciCode | 51.2% | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated) |
| SWE Atlas - Codebase QnA | 48.1% | DefaultGLM 5.2 (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 2 (Scale rank accounts for CI); ±5.07; entry added 2026-06-23 |
| SWE Atlas - Refactoring | 42.4% | DefaultGLM 5.2 (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-06-23 |
| SWE Atlas - Test Writing | 41.5% | DefaultGLM 5.2 (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 2 (Scale rank accounts for CI); ±5.96; entry added 2026-06-23 |
| SWE Marathon | 13.0% | Maxmax effort | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | Run by Abundant AI at max effort (pre-v1.1) |
| SWE-Bench Pro (public, v1) | 62.1% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | OpenHands | tailored instruction prompt, 400K ctx |
| SWE-bench Verified | 82.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.689 stderr; $0.710/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| SWE-bench Verified | 78.7% | Maxglm-5.2_max | Independent testIndependentEpoch AI ↗ | 25 Jun 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| SWE-rebench | 62.9% | HighGLM-5.2 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | 1 Oct 2026 | SWE-rebench standard scaffold | time window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.19; pass@5 81.1%; $1.40/problem |
| tau2-bench | 99.1% | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 77.9% | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 81.0% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | Terminus 2 | 82.7 with Claude Code 2.1.167 ('best reported harness', avg of 5 runs) |
| Terminal-Bench 3.0 | 4.6% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 3.3B, run cost $3.4k |
| Terminal-Bench 3.0 | 4.6% | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | GLM-5.2 column in the GLM-5.3 launch table |
| Terminal-Bench 4.0 | 1.0% | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated) |
| Toolathlon | 48.2% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | 16 Jun 2026 | — | Reported as 'Tool-Decathlon' |
| Toolathlon-Verified | 59.9% | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | GLM-5.2 column in the GLM-5.3 launch table |
| Vals SRE Bench | 0.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±0 stderr; $26.575/test |
| Vending-Bench 2 | $8,314 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$1,084; rank 10; only top 10 rendered server-side (57 more behind 'Show more') |
| Z.ai Code Bench (internal) | 23.4% | MaxMax effort | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | ~96K output tokens/task (GLM-5.3 blog text) |