Skip to content
Bencher

Models · Z.ai (Zhipu) · Out sinceReleased 26 Aug 2026

GLM-5.3-Flash#10 for research and analysis.#10 for research and analysis, best at its default setting.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

GLM-5.3-Flash is made by Z.ai (Zhipu). Among the models we track it ranks #10 for research and analysis. It's cheap to use. Its makers have published it, so you can run it on your own servers.

320B total / 18B active, hybrid sparse+linear attention, first natively multimodal GLM-5 model, MIT license. reasoning_effort low/high/max (default max). Cached input $0.03.

Writing & creativity
—
Research & analysis
74.6 / 100 · #10
Coding
—
Price
Cheap$0.15 / $0.50
Price per 1M (blended)Blended / 1M
$0.24
MemoryContext
1M
Longest answerMax output
—
Test resultsResults
35 (26 independent26 indep.)
Out sinceReleased
26 Aug 2026
Made byVendor
Z.ai (Zhipu)
UnderstandsInputs
text, image, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowHighMax

Highlighted: where it did best for research and analysis (its standard thinking level).

Highlighted: dominant setting in its research and analysis composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as z-ai/glm-5.3-flash.

Route it as z-ai/glm-5.3-flash at $0.15 in / $0.50 out per 1M tokens, 1M context. Listed since 26 Aug 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

LicenceLicence
MIT
OK for business useCommercial use
Yes
SizeParams (total / active)
320B / 18B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Needs one full AI server.

≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 336 GB at 8-bit, 176 GB at 4-bit (+10–30% for KV cache). MoE: 18B active per token.

Languages: Languages: HF metadata: en, zh. EuroEval Swedish rank 1 among open models (rank score 1.26).

Training it further: Fine-tuning: Card links Unsloth guide; no vendor SFT recipe; newly trained base model not released.

Quantisations: fp8 (default release), bf16 (zai-org/GLM-5.3-Flash-BF16), nvfp4 (NVIDIA: nvidia/GLM-5.3-Flash-NVFP4), gguf (community: Unsloth). Engines: SGLang, vLLM, TokenSpeed, Transformers, KTransformers, llama.cpp (via Unsloth guide).

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗github.com ↗

Thinking level

Reasoning effort

How long should GLM-5.3-FlashWhere GLM-5.3-Flashthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

CursorBench

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11452DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR80.0%DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy27.5%DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate27.6%DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index7.5DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam26.3%Maxreasoning_effort=max (default)Maker's own figureVendor-reportedZ.ai ↗26 Aug 2026Claude Code
Artificial Analysis Intelligence Index41.8DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug glm-5-3-flash; list price $0.15/0.5 per 1M in/out; cost to run AA Intelligence Index $0.25/task
Artificial Analysis output speed45 tok/sDefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 3.3s; list price $0.15/0.5 per 1M in/out
AutomationBench48.8%Maxreasoning_effort=max (default)Maker's own figureVendor-reportedZ.ai ↗26 Aug 2026—v1.0.6
CursorBench26.9%LowGLM 5.3 Flash LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 57; $0.15/task; 17,831 tokens/task; 58 steps/task
CursorBench31.1%HighGLM 5.3 Flash HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 49; $0.25/task; 35,104 tokens/task; 84 steps/task
CursorBench36.8%MaxGLM 5.3 Flash MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 34; $0.39/task; 56,410 tokens/task; 118 steps/task
DeepSWE63.4%Maxglm-5.3-flash_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.0%; ±4.4; 4 runs; $0.48/task
DeepSWE63.4%Maxreasoning_effort=max (default)Maker's own figureVendor-reportedZ.ai ↗26 Aug 2026mini-swe-agentDeepSWE v1.1
EuroEval Swedish (generative)1.3Defaultzai-org/GLM-5.3-Flash (val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 1; ±0.12; lower is better; task scores (first metric): SweDN summarisation 41.72 ± 0.43, Skolprov 78.21 ± 3.29, Swedish facts 61.66 ± 3.18, ScaLA-sv 77.72 ± 0.88
FrontierCode31.8%Maxglm-5.3-flash_maxIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026chiselread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5
GDPval-AA v21773Maxreasoning_effort=max (default)Maker's own figureVendor-reportedZ.ai ↗26 Aug 2026—Evaluated by Artificial Analysis
GDPval-AA v2.11641DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1618.06-1663.5
GPQA Diamond91.2%DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-3-flash)
Humanity's Last Exam39.9%DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-3-flash)
Humanity's Last Exam (with tools)55.3%Maxreasoning_effort=max (default)Maker's own figureVendor-reportedZ.ai ↗26 Aug 2026—full set
IOI (Vals)52.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±2.619 stderr; $0.429/test
LegalBench (Vals)83.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id zai/glm-5.3-flash; rank 43/149; ±0.427 stderr; $0.000278/test
LMArena Code Arena (WebDev)1615Defaultglm-5.3-flashIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 24 (CI rank 17-25); 95% CI 1607-1623; 9648 votes
LMArena Text - Coding category1523Defaultglm-5.3-flashIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 26 (CI rank 8-62); 95% CI 1515-1532; 5833 votes
NL2Repo-Bench56.3%Maxreasoning_effort=max (default)Maker's own figureVendor-reportedZ.ai ↗26 Aug 2026—1M ctx
SciCode51.6%DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-3-flash)
SWE-bench Verified92.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.214 stderr; $0.018/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.184.3%Maxreasoning_effort=max (default)Maker's own figureVendor-reportedZ.ai ↗26 Aug 2026Claude Code 2.1.2076h timeout
Terminal-Bench 2.184.3%DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-3-flash)
Terminal-Bench 4.032.8%DefaultGLM 5.3 FlashIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug glm-5-3-flash)
Toolathlon-Verified78.4%Maxreasoning_effort=max (default)Maker's own figureVendor-reportedZ.ai ↗26 Aug 2026—
Vals Legal Research Bench45.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id zai/glm-5.3-flash; rank 14/72; ±3.459 stderr; $0.120196/test
Vals TaxEval v275.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id zai/glm-5.3-flash; rank 16/145; ±0.836 stderr; $0.001505/test
Z.ai Code Bench (internal)29.0%MaxMax effortMaker's own figureVendor-reportedZ.ai ↗26 Aug 2026—Z.ai Code Bench v1.0 (blog text)