Models · Z.ai (Zhipu) · Out sinceReleased 14 Aug 2026
GLM-5.3#13 for writing.#13 for writing, best at max effort.
GLM-5.3 is made by Z.ai (Zhipu). Among the models we track it ranks #13 for writing, #13 for coding, #19 for research and analysis. It's mid-priced to use. Its makers have published it, so you can run it on your own servers. GLM-5.3 is the strongest freely downloadable model that still fits on a single AI server. Its licence lets you use it commercially, so your documents never have to leave your own or an EU data centre.
Same base as GLM-5.2, post-training only. Thinking cannot be disabled; reasoning_effort low/high/max (default max; max recommended for coding). Weights released ~2 weeks after launch (license 'glm-5.3'). Cached input $0.26.
- 56.6 / 100 · #13
- 64.8 / 100 · #19
- 55.5 / 100 · #13
- Mid-priced$1.40 / $4.40
- $2.15
- 1M
- 131K
- 85 (68 independent68 indep.)
- 14 Aug 2026
- Z.ai (Zhipu)
- text
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (Maximum thinking).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as z-ai/glm-5.3.
Route it as z-ai/glm-5.3 at $0.30 in / $4.40 out per 1M tokens, 1M context. Listed since 18 Aug 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 744B / 40B
- No
Needs one full AI server.
≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 781 GB at 8-bit, 409 GB at 4-bit (+10–30% for KV cache). MoE: 40B active per token.
Licence conditions: Keep copyright/permission notice; comply with applicable law. Only MaaS businesses with >US$10B revenue over 12 months must pass a Z.AI security review before commercial use.
Restrictions: Keep copyright/permission notice; comply with applicable law. Only MaaS businesses with >US$10B revenue over 12 months must pass a Z.AI security review before commercial use.
Languages: Languages: HF metadata: en, zh.
Training it further: Fine-tuning: Model card links an Unsloth guide (unsloth.ai/docs/models/GLM-5.3); no vendor SFT recipe. Same base as GLM-5.2, base checkpoint not released.
Quantisations: fp8 (default release), bf16 (zai-org/GLM-5.3-BF16), nvfp4 (NVIDIA: nvidia/GLM-5.3-NVFP4), gguf (community: Unsloth dynamic). Engines: SGLang, vLLM, TokenSpeed, Transformers, KTransformers, llama.cpp (via Unsloth GGUF guide), vLLM-Ascend/xLLM (Ascend NPU).
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗github.com ↗
Thinking level
Reasoning effort
How long should GLM-5.3Where GLM-5.3think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
SciCode
Best at MaxBest at Max
Artificial Analysis Intelligence Index
Best at MaxBest at Max
AA-LCR
Best at MaxBest at Max
AA-Omniscience Accuracy
Best at Low, worse abovePeaks at Low
AA-Omniscience Hallucination Rate
Best at MaxBest at Max
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Z.ai Code Bench (internal)
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1511 | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 72.7% | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 79.7% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 33.9% | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 33.9% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 64.0% | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 29.6% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | -8.4 | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 14.3 | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 28.5% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | Claude Code | ALE-CLI, 105 tasks |
| Artificial Analysis Coding Agent Index | 53.6 | DefaultGLM-5.3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Opencode | agent Opencode; components: DeepSWE v1.1 61.4, SWE-Atlas-QnA 59.4, Terminal-Bench v4 39.9; avg cost $4.24/task; avg wall time 48 min/task |
| Artificial Analysis Intelligence Index | 34.3 | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug glm-5-3-low; list price $1.4/4.4 per 1M in/out; cost to run AA Intelligence Index $0.85/task |
| Artificial Analysis Intelligence Index | 44.8 | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug glm-5-3; list price $1.4/4.4 per 1M in/out; cost to run AA Intelligence Index $2.01/task |
| Artificial Analysis output speed | 69 tok/s | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 3.1s; list price $1.4/4.4 per 1M in/out |
| Artificial Analysis output speed | 65 tok/s | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 3.3s; list price $1.4/4.4 per 1M in/out |
| AutomationBench | 48.2% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | AutomationBench v1.0.6 |
| CursorBench | 33.3% | LowGLM 5.3 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 43; $2.04/task; 31,983 tokens/task; 81 steps/task |
| CursorBench | 38.0% | HighGLM 5.3 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 30; $3.24/task; 60,031 tokens/task; 114 steps/task |
| CursorBench | 42.6% | MaxGLM 5.3 Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 20; $5.05/task; 96,387 tokens/task; 166 steps/task |
| CyberGym | 84.5% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | Claude Code 2.1.207 | single-run pass@1 over 1,507 tasks |
| DeepSWE | 69.0% | Maxglm-5.3_max | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 87.6%; ±3.0; 4 runs; $3.99/task |
| DeepSWE | 66.9% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | mini-swe-agent | DeepSWE v1.1; temp 0.95, 6h timeout, 400K ctx |
| DeepSWE | 61.4% | DefaultGLM-5.3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Opencode | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| Epoch Capabilities Index | 155.8 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 153.7-158.3; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 2075 | DefaultGLM-5.3 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 6; rubric score 17.04/20; slop 8.42; avg length 5913 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| ExploitBench | 54.4% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | Claude Code 2.1.207 | |
| FrontierCode | 40.1% | Maxglm-5.3_max | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | chisel | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5 |
| FrontierMath (Tiers 1-3) | 68.8% | Maxglm-5.3_max | Independent testIndependentEpoch AI ↗ | 25 Aug 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierSWE | 78.1% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | Run by Proximal; dominance score as of 2026-08-14 |
| FrontierSWE | 30.2% | Default | Independent testIndependentFrontierSWE ↗ | 1 Oct 2026 | proximus | FrontierSWE V2, mean@5 over 34 tasks (20h budget); ±11.5; $97.22/trial; 17.0h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort |
| GDPval-AA v2 | 1769 | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | Evaluated by Artificial Analysis |
| GDPval-AA v2.1 | 1644 | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1621.04-1666.21 |
| GPQA Diamond | 91.7% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3) |
| HLE Diamond | 16.4% | Defaultglm-5.3 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 11 (Scale rank accounts for CI); ±2.3; entry added 2025-11-06 |
| Humanity's Last Exam | 36.6% | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3-low) |
| Humanity's Last Exam | 42.3% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3) |
| Humanity's Last Exam (with tools) | 62.5% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | 300K ctx with context management |
| IOI (Vals) | 68.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±7.476 stderr; $7.671/test |
| LegalBench (Vals) | 84.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id zai/glm-5.3; rank 27/149; ±0.4 stderr; $0.006513/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 3 | MaxGLM-5.3 (max) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 6/56; Thurstone comparison score (centered at 0); est. win chance 88%; 95% bootstrap 2.928 to 3.127 |
| LMArena Code Arena (WebDev) | 1622 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 19 (CI rank 14-24); 95% CI 1614-1630; 7457 votes |
| LMArena Text - Creative Writing | 1455 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 31 (rank range 13-61); 95% CI 1444.8-1465.7; 3895 votes; style-controlled |
| LMArena Text - Expert | 1519 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 4-54); 95% CI 1504.0-1532.9; 1753 votes; style-controlled |
| LMArena Text - Hard Prompts | 1505 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 21 (rank range 10-41); 95% CI 1498.7-1511.9; 11072 votes; style-controlled |
| LMArena Text - Instruction Following | 1478 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 21 (rank range 10-43); 95% CI 1469.5-1486.2; 5877 votes; style-controlled |
| LMArena Text - Longer Query | 1489 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 27 (rank range 13-54); 95% CI 1481.1-1496.2; 8106 votes; style-controlled |
| LMArena Text - Multi-Turn | 1481 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 39 (rank range 8-76); 95% CI 1468.5-1492.7; 2623 votes; style-controlled |
| LMArena Text - Non-English | 1464 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 31 (rank range 18-50); 95% CI 1457.7-1471.1; 10284 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1475 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 32 (rank range 9-65); 95% CI 1464.3-1486.0; 3253 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1473 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 51 (rank range 12-102); 95% CI 1457.2-1488.4; 1520 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1469 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 24 (rank range 9-46); 95% CI 1460.1-1478.5; 4954 votes; style-controlled |
| LMArena Text (overall) | 1479 | Maxglm-5.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 26 (CI rank 13-44); 95% CI 1474-1485; 17268 votes |
| MCP Atlas | 84.2% | DefaultGLM 5.3 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±2.15; entry added 2026-09-17 |
| NL2Repo-Bench | 58.0% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | 1M ctx, anti-hacking judge |
| PostTrainBench v1.1 | 39.8% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | Claude Code 2.1.207 | weighted avg of 3 runs |
| ProgramBench (Almost Solved) | 19.0% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | ProgramBench 'Almost Solved' |
| ProgramBench (fully resolved) | 1.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.862 stderr; $21.957/test; strict fully-resolved rate |
| SciCode | 42.0% | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3-low) |
| SciCode | 59.0% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3) |
| SimpleBench | 66.2% | DefaultGLM 5.3 | Independent testIndependentSimpleBench ↗ | 19 Aug 2026 | — | AVG@5, temp 0.7; rank 22nd; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 41.0% | Maxglm-5.3_max | Independent testIndependentEpoch AI ↗ | 28 Aug 2026 | — | Epoch-run (no tools); ±1.56 stderr |
| SkillsBench | 47.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.522 stderr; $0.714/test |
| SWE Atlas - Codebase QnA | 59.4% | DefaultGLM-5.3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Opencode | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Marathon | 42.5% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | Claude Code 2.1.207 | SWE-Marathon v1.1; some anti-cheat checks replaced by LLM inspection |
| SWE-Bench Pro V2 (full) | 95.6% | MaxGLM-5.3 (mini-swe-agent) max | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 5 (Scale rank accounts for CI); ±1.3; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-Bench Pro V2 (hard) | 84.3% | MaxGLM-5.3 (mini-swe-agent) max | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 7 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-bench Verified | 95.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±0.938 stderr; $0.338/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 83.9% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3) |
| Terminal-Bench 2.1 | 88.2% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | Claude Code 2.1.207 | temp 1.0, top_p 1, max_new_tokens 65,536, 6h timeout |
| Terminal-Bench 3.0 | 32.4% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 5.6B, run cost $1.8k |
| Terminal-Bench 3.0 | 28.3% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | Claude Code 2.1.207 | avg@3, 400K ctx, 600 turns, 10h timeout, official verifiers |
| Terminal-Bench 4.0 | 34.8% | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3-low) |
| Terminal-Bench 4.0 | 41.9% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3) |
| Terminal-Bench 4.0 | 41.8% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.23 95% CI; 330 trials; total run cost $2728; model release 2026-08-14 |
| Terminal-Bench 4.0 | 38.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.821 stderr; $9.374/test |
| Terminal-Bench 4.0 | 39.9% | DefaultGLM-5.3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Opencode | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Toolathlon-Verified | 73.0% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | official evaluation service, avg of 3 runs |
| Vals Code Migration | 44.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.289 stderr; $24.910/test |
| Vals Index | 53.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.305 stderr; $7.249/test |
| Vals Legal Research Bench | 49.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id zai/glm-5.3; rank 7/72; ±3.475 stderr; $2.241388/test |
| Vals Public Benefits Bench v1.1 | 68.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id zai/glm-5.3; rank 8/45; ±1.208 stderr; $0.946832/test |
| Vals TaxEval v2 | 72.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id zai/glm-5.3; rank 66/145; ±0.879 stderr; $0.032/test |
| Vals Vibe Code Bench | 78.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±3.94 stderr; $12.450/test |
| Z.ai Code Bench (internal) | 31.4% | HighHigh effort | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | ~50K output tokens/task (blog text) |
| Z.ai Code Bench (internal) | 34.5% | MaxMax effort | Maker's own figureVendor-reportedZ.ai ↗ | 14 Aug 2026 | — | Z.ai Code Bench end-to-end completion; ~75K output tokens/task (read from blog text) |