Models · Z.ai (Zhipu) · Out sinceReleased 26 Aug 2026
GLM-5.3-Flash#10 for research and analysis.#10 for research and analysis, best at its default setting.
GLM-5.3-Flash is made by Z.ai (Zhipu). Among the models we track it ranks #10 for research and analysis. It's cheap to use. Its makers have published it, so you can run it on your own servers.
320B total / 18B active, hybrid sparse+linear attention, first natively multimodal GLM-5 model, MIT license. reasoning_effort low/high/max (default max). Cached input $0.03.
- —
- 74.6 / 100 · #10
- —
- Cheap$0.15 / $0.50
- $0.24
- 1M
- —
- 35 (26 independent26 indep.)
- 26 Aug 2026
- Z.ai (Zhipu)
- text, image, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (its standard thinking level).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as z-ai/glm-5.3-flash.
Route it as z-ai/glm-5.3-flash at $0.15 in / $0.50 out per 1M tokens, 1M context. Listed since 26 Aug 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 320B / 18B
- No
Needs one full AI server.
≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 336 GB at 8-bit, 176 GB at 4-bit (+10–30% for KV cache). MoE: 18B active per token.
Languages: Languages: HF metadata: en, zh. EuroEval Swedish rank 1 among open models (rank score 1.26).
Training it further: Fine-tuning: Card links Unsloth guide; no vendor SFT recipe; newly trained base model not released.
Quantisations: fp8 (default release), bf16 (zai-org/GLM-5.3-Flash-BF16), nvfp4 (NVIDIA: nvidia/GLM-5.3-Flash-NVFP4), gguf (community: Unsloth). Engines: SGLang, vLLM, TokenSpeed, Transformers, KTransformers, llama.cpp (via Unsloth guide).
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗github.com ↗
Thinking level
Reasoning effort
How long should GLM-5.3-FlashWhere GLM-5.3-Flashthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
CursorBench
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1452 | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 80.0% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 27.5% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 27.6% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 7.5 | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 26.3% | Maxreasoning_effort=max (default) | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | Claude Code | |
| Artificial Analysis Intelligence Index | 41.8 | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug glm-5-3-flash; list price $0.15/0.5 per 1M in/out; cost to run AA Intelligence Index $0.25/task |
| Artificial Analysis output speed | 45 tok/s | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 3.3s; list price $0.15/0.5 per 1M in/out |
| AutomationBench | 48.8% | Maxreasoning_effort=max (default) | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | — | v1.0.6 |
| CursorBench | 26.9% | LowGLM 5.3 Flash Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 57; $0.15/task; 17,831 tokens/task; 58 steps/task |
| CursorBench | 31.1% | HighGLM 5.3 Flash High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 49; $0.25/task; 35,104 tokens/task; 84 steps/task |
| CursorBench | 36.8% | MaxGLM 5.3 Flash Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 34; $0.39/task; 56,410 tokens/task; 118 steps/task |
| DeepSWE | 63.4% | Maxglm-5.3-flash_max | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.0%; ±4.4; 4 runs; $0.48/task |
| DeepSWE | 63.4% | Maxreasoning_effort=max (default) | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | mini-swe-agent | DeepSWE v1.1 |
| EuroEval Swedish (generative) | 1.3 | Defaultzai-org/GLM-5.3-Flash (val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 1; ±0.12; lower is better; task scores (first metric): SweDN summarisation 41.72 ± 0.43, Skolprov 78.21 ± 3.29, Swedish facts 61.66 ± 3.18, ScaLA-sv 77.72 ± 0.88 |
| FrontierCode | 31.8% | Maxglm-5.3-flash_max | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | chisel | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5 |
| GDPval-AA v2 | 1773 | Maxreasoning_effort=max (default) | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | — | Evaluated by Artificial Analysis |
| GDPval-AA v2.1 | 1641 | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1618.06-1663.5 |
| GPQA Diamond | 91.2% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3-flash) |
| Humanity's Last Exam | 39.9% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3-flash) |
| Humanity's Last Exam (with tools) | 55.3% | Maxreasoning_effort=max (default) | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | — | full set |
| IOI (Vals) | 52.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.619 stderr; $0.429/test |
| LegalBench (Vals) | 83.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id zai/glm-5.3-flash; rank 43/149; ±0.427 stderr; $0.000278/test |
| LMArena Code Arena (WebDev) | 1615 | Defaultglm-5.3-flash | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 24 (CI rank 17-25); 95% CI 1607-1623; 9648 votes |
| LMArena Text - Coding category | 1523 | Defaultglm-5.3-flash | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 26 (CI rank 8-62); 95% CI 1515-1532; 5833 votes |
| NL2Repo-Bench | 56.3% | Maxreasoning_effort=max (default) | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | — | 1M ctx |
| SciCode | 51.6% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3-flash) |
| SWE-bench Verified | 92.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.214 stderr; $0.018/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 84.3% | Maxreasoning_effort=max (default) | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | Claude Code 2.1.207 | 6h timeout |
| Terminal-Bench 2.1 | 84.3% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3-flash) |
| Terminal-Bench 4.0 | 32.8% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug glm-5-3-flash) |
| Toolathlon-Verified | 78.4% | Maxreasoning_effort=max (default) | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | — | |
| Vals Legal Research Bench | 45.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id zai/glm-5.3-flash; rank 14/72; ±3.459 stderr; $0.120196/test |
| Vals TaxEval v2 | 75.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id zai/glm-5.3-flash; rank 16/145; ±0.836 stderr; $0.001505/test |
| Z.ai Code Bench (internal) | 29.0% | MaxMax effort | Maker's own figureVendor-reportedZ.ai ↗ | 26 Aug 2026 | — | Z.ai Code Bench v1.0 (blog text) |