Skip to content
Bencher

Models · Z.ai (Zhipu) · Out sinceReleased 16 Jun 2026

GLM-5.2#20 for coding.#20 for coding, best at its default setting.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

GLM-5.2 is made by Z.ai (Zhipu). Among the models we track it ranks #20 for coding. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.

MIT license, 1M context, IndexShare sparse attention. Introduced effort levels (High/Max exposed in Coding Plan).

Writing & creativity
—
Research & analysis
—
Coding
49.8 / 100 · #20
Price
Mid-priced$1.40 / $4.40
Price per 1M (blended)Blended / 1M
$2.15
MemoryContext
1M
Longest answerMax output
131K
Test resultsResults
51 (26 independent26 indep.)
Out sinceReleased
16 Jun 2026
Made byVendor
Z.ai (Zhipu)
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

HighMax

Highlighted: where it did best for coding (its standard thinking level).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as z-ai/glm-5.2.

Route it as z-ai/glm-5.2 at $1.40 in / $4.40 out per 1M tokens, 1M context. Listed since 16 Jun 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

Details comingselfHost data pending

This model can be downloaded and run on your own servers. We're collecting its licence, size and hardware needs now.Open-weights model; licence, parameter count, hardware tier and base-model availability will appear after the next data run.

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Agents' Last Exam23.8%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗14 Aug 2026—GLM-5.2 column in the GLM-5.3 launch table; ALE-CLI
AIME 202699.2%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—
Artificial Analysis Intelligence Index33.7MaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug glm-5-2; list price $1.4/4.4 per 1M in/out; cost to run AA Intelligence Index $1.47/task (AA marks this variant deprecated)
Artificial Analysis output speed82 tok/sMaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 3.1s; list price $1.4/4.4 per 1M in/out (AA marks this variant deprecated)
AutomationBench26.2%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗14 Aug 2026—GLM-5.2 column in the GLM-5.3 launch table; v1.0.6
CyberGym77.2%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗14 Aug 2026—GLM-5.2 column in the GLM-5.3 launch table
DeepSWE46.2%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026mini-swe-agentofficial pier framework, 2h timeout
Design Arena (all categories)1297Defaultglm-5.2Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±1.9 SE; 42958 battles; win rate 52.1%
Design Arena (fullstack)1239Defaultglm-5.2Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±7.3 SE; 2758 battles; win rate 60.5%
EQ-Bench Creative Writing v3 (Elo)1757Defaultzai-org/GLM-5.2Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 29; rubric score 16.44/20; slop 13.11; avg length 6104 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
ExploitBench24.4%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗14 Aug 2026—GLM-5.2 column in the GLM-5.3 launch table
FrontierSWE74.4%Maxmax effortMaker's own figureVendor-reportedZ.ai ↗16 Jun 2026—Run by Proximal at max effort; dominance as of 2026-06-16
GDPval-AA v21508Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗14 Aug 2026—GLM-5.2 column in the GLM-5.3 launch table
GPQA Diamond89.5%MaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated)
GPQA Diamond91.2%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—
HMMT February 202692.5%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—HMMT Feb. 2026
Humanity's Last Exam41.1%MaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated)
Humanity's Last Exam40.5%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—Text-only subset
Humanity's Last Exam (with tools)54.7%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—300K ctx, no context management
IMO-AnswerBench91.0%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—
LLM Creative Story-Writing Benchmark (Lech Mazur)0.5MaxGLM-5.2 (max)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 22/56; Thurstone comparison score (centered at 0); est. win chance 57%; 95% bootstrap 0.380 to 0.561
LMArena Code Arena (WebDev)1605Maxglm-5.2-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 25 (CI rank 22-26); 95% CI 1598-1611; 14543 votes
MCP Atlas77.8%Defaultglm-5p2Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±2.6; entry added 2026-07-20
MCP Atlas76.8%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—Public set, think mode
NL2Repo-Bench48.9%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—400K ctx
PostTrainBench v1.134.3%Maxmax effortMaker's own figureVendor-reportedZ.ai ↗16 Jun 2026—Run by PostTrainBench at max effort
ProgramBench (Almost Solved)9.5%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗14 Aug 2026—GLM-5.2 column in the GLM-5.3 launch table; ProgramBench 'Almost Solved'
ProgramBench (avg test pass rate)64.6%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 8.5%; avg cost $25.36/task
ProgramBench (fully resolved)0.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.5 stderr; $12.577/test; strict fully-resolved rate
ProgramBench (fully resolved)63.7%Maxmax effortMaker's own figureVendor-reportedZ.ai ↗16 Jun 2026Claude Code 2.1.156200 instances, reasoning_effort=max
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 8.5%; avg cost $25.36/task
SciCode51.2%MaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated)
SWE Atlas - Codebase QnA48.1%DefaultGLM 5.2 (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 2 (Scale rank accounts for CI); ±5.07; entry added 2026-06-23
SWE Atlas - Refactoring42.4%DefaultGLM 5.2 (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-06-23
SWE Atlas - Test Writing41.5%DefaultGLM 5.2 (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 2 (Scale rank accounts for CI); ±5.96; entry added 2026-06-23
SWE Marathon13.0%Maxmax effortMaker's own figureVendor-reportedZ.ai ↗16 Jun 2026—Run by Abundant AI at max effort (pre-v1.1)
SWE-Bench Pro (public, v1)62.1%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026OpenHandstailored instruction prompt, 400K ctx
SWE-bench Verified82.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.689 stderr; $0.710/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
SWE-bench Verified78.7%Maxglm-5.2_maxIndependent testIndependentEpoch AI ↗25 Jun 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)
SWE-rebench62.9%HighGLM-5.2 [high]Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.19; pass@5 81.1%; $1.40/problem
tau2-bench99.1%MaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated)
Terminal-Bench 2.177.9%MaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated)
Terminal-Bench 2.181.0%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026Terminus 282.7 with Claude Code 2.1.167 ('best reported harness', avg of 5 runs)
Terminal-Bench 3.04.6%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 3.3B, run cost $3.4k
Terminal-Bench 3.04.6%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗14 Aug 2026—GLM-5.2 column in the GLM-5.3 launch table
Terminal-Bench 4.01.0%MaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-2) (AA marks this variant deprecated)
Toolathlon48.2%Defaultthinking (effort not stated)Maker's own figureVendor-reportedZ.ai ↗16 Jun 2026—Reported as 'Tool-Decathlon'
Toolathlon-Verified59.9%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗14 Aug 2026—GLM-5.2 column in the GLM-5.3 launch table
Vals SRE Bench0.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±0 stderr; $26.575/test
Vending-Bench 2$8,314DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$1,084; rank 10; only top 10 rendered server-side (57 more behind 'Show more')
Z.ai Code Bench (internal)23.4%MaxMax effortMaker's own figureVendor-reportedZ.ai ↗14 Aug 2026—~96K output tokens/task (GLM-5.3 blog text)