GLM 5: Every Version from GLM-5 to GLM-5.3-Flash, Priced and Benchmarked

GLM 5 is Z.ai's open-weight coding model family: GLM-5 (Feb 2026), GLM-5.1 (Apr), GLM-5.2 (Jun, 1M context), GLM-5.3 (Aug, Intelligence Index v4.3 score of 45), and GLM-5.3-Flash (320B/18B). Release dates, parameters, licenses, Z.ai vs Morph pricing (GLM-5.3 at $1.00/$3.41 per M on Morph), the Claude Opus comparison with cost per task, the RAM and VRAM each quantization needs to run it locally, provider throughput, and how to call it.

September 1, 2026 · 2 min read
Feb 11 to Aug 26, 2026
5 releases
Feb 11 to Aug 26, 2026
Total / active params, GLM-5 to GLM-5.3
753B / 40B
Total / active params, GLM-5 to GLM-5.3
GLM-5.3 Intelligence Index v4.3 (Opus 4.8: 42)
45
GLM-5.3 Intelligence Index v4.3 (Opus 4.8: 42)
GLM-5.3 on Morph, per 1M tokens
$1.00 / $3.41
GLM-5.3 on Morph, per 1M tokens

TL;DR

GLM 5 is Z.ai's family of open-weight coding models, released across 2026 as GLM-5, GLM-5.1, GLM-5.2, GLM-5.3, and GLM-5.3-Flash. The large versions are 753B mixture-of-experts models with about 40B active parameters and a 1M context since GLM-5.2. GLM-5.3 is the current best, scoring 45 on the independent Artificial Analysis Intelligence Index v4.3 against Claude Opus 4.8's 42, checked September 7, 2026, at $1.40/$4.40 per million tokens.

Last updated September 7, 2026. Morph rates come from the live billing table; every external number links to its source at the bottom of the page.

28 to 45
Artificial Analysis Intelligence Index v4.3, GLM-5 (February) to GLM-5.3 (August). Same 753B base, same $1.40/$4.40 list price since April. Claude Opus 4.8 scores 42 at $5/$25.
Artificial Analysis model pages, Intelligence Index v4.3, checked September 7, 2026

GLM 5 (often written GLM5) is not one model. It is the name Z.ai (Zhipu AI) has used for five releases in 2026: GLM-5.1, GLM-5.2, GLM-5.3, and GLM-5.3-Flash all sit on top of the original GLM-5 from February 11. The four large models share one 753B mixture-of-experts shape with about 40B active parameters per token. Context grew from 200K to 1M at GLM-5.2. Capability grew every release without a new pretraining run after GLM-5.2: GLM-5.3 is scaled post-training on the GLM-5.2 base.

Where it stands against Claude: on the independent Intelligence Index v4.3, checked September 7, 2026, GLM-5.3 scores 45, Claude Opus 4.8 scores 42, Opus 5 scores 51. Z.ai's own table has GLM-5.3 ahead of Opus 4.8 on 12 of 16 benchmarks. Z.ai lists it at $1.40/$4.40 per million tokens against Opus's $5/$25. Morph serves GLM-5.3 as morph-glm53-744b at $1.00/$3.41 and GLM-5.3-Flash as morph-glm53flash at $0.10/$0.35, both with the full 1M context.

What GLM 5 is

A 753B MoE (about 40B active) with DeepSeek Sparse Attention, trained on 28.5T tokens, released as open weights. Four point releases in seven months, each a post-training or context upgrade on the same base. GLM-5.3-Flash adds a 320B/18B variant with vision.

The catch

Reasoning is always on and defaults to max effort, so output tokens run high: GLM-5.3 emitted 210M tokens across the Artificial Analysis suite against a 120M open-weights median. Cheap per token, less cheap per task. Cap effort at high for routine work.

GLM 5 Quick Reference

Everything on one screen for the current release. Older versions are in the family table below.

GLM-5.3 and GLM-5.3-Flash, September 7, 2026
GLM-5.3GLM-5.3-Flash
DeveloperZ.ai (formerly Zhipu AI), BeijingZ.ai
ReleasedAugust 14, 2026 (weights August 25)August 26, 2026
Parameters (total / active)753B / 40B320B / 18B
Context / max output1M / 128K1M / 128K
LicenseGLM-5.3 License (MIT-like, $10B-revenue clause)MIT
Weightshuggingface.co/zai-org/GLM-5.3huggingface.co/zai-org/GLM-5.3-Flash
Vision inputNoYes
Intelligence Index v4.3 (AA)4542
Z.ai list, in / out per M$1.40 / $4.40$0.15 / $0.50
Morph, in / cached / out per M$1.00 / $0.20 / $3.41$0.10 / $0.02 / $0.35
Morph model idmorph-glm53-744bmorph-glm53flash
Subscription optionGLM Coding Plan, from $18/monthGLM Coding Plan, from $18/month
Smallest useful local quantUD-IQ2_M, 238.6 GB, needs 245 GB RAM+VRAMUD-Q2_K_XL, 108.72 GB, needs 115 GB
GPU serving recipe8x H200/H20 FP8, 8x B200 for full 1M4x B200, via the vLLM docker image
ReasoningAlways on, low/high/max effortAlways on, low/high/max effort

Local quant sizes and memory floors are Unsloth's published GGUF tables. GPU recipes are the official vLLM recipes. Everything else is linked in Sources.

GLM 5 Family at a Glance

Every row below is sourced. Z.ai prices are list rates from its pricing docs. Morph prices are the live billing rates for the two versions Morph serves. Benchmarks in the last column are Z.ai-reported unless marked AA (Artificial Analysis, independent).

GLM 5 versions, September 2026
VersionReleasedParams (total / active)ContextWeightsZ.ai in / out per MMorph per MHeadline number
GLM-5Feb 11, 2026744B / 40B200KMIT, zai-org/GLM-5$1.00 / $3.20not servedSWE-bench Verified 77.8; AA Index v4.3 28
GLM-5.1Apr 7, 2026744B / 40B200KMIT, zai-org/GLM-5.1$1.40 / $4.40not servedSWE-bench Pro 58.4; AA Index v4.3 27
GLM-5.2Jun 13, 2026753B / ~40B1MMIT, zai-org/GLM-5.2$1.40 / $4.40alias of morph-glm53-744bSWE-bench Pro 62.1; AA Index v4.3 39
GLM-5.3Aug 14, 2026 (weights Aug 25)753B / 40B1MGLM-5.3 License, zai-org/GLM-5.3$1.40 / $4.40$1.00 / $3.41Terminal-Bench 2.1 88.2; AA Index v4.3 45
GLM-5.3-FlashAug 26, 2026320B / 18B1MMIT, zai-org/GLM-5.3-Flash$0.15 / $0.50$0.10 / $0.35AA Index v4.3 42; vision

Parameter counts: Z.ai's docs say 744B for GLM-5 and GLM-5.1, while the Hugging Face checkpoint metadata for both reports 753.9B, the figure Z.ai rounds to 753B for GLM-5.2 and GLM-5.3. The config files for GLM-5 and GLM-5.1 are identical (78 layers, 256 routed experts plus 1 shared, 8 active per token). GLM-5.3-Flash pricing: Z.ai's pricing page shows a promotional $0.075 / $0.25; Artificial Analysis lists $0.15 / $0.50. Morph's morph-glm52-744b id still resolves and routes to the GLM-5.3 stack.

Which one to pick

GLM-5.3 for anything hard: it is the top version at the same list price as 5.1 and 5.2. GLM-5.3-Flash for volume and vision at about a tenth of the price. GLM-5.2 only if the large model must be MIT-licensed. GLM-5 and GLM-5.1 are superseded: 200K context, Index 41, and (for GLM-5.1) the same price as 5.3.

What Changed at Each Version

GLM-5 (February 11, 2026): the base

The only new pretraining run in the family. A 744B-total MoE with 40B active per token, 256 routed experts plus 1 shared with 8 active, and DeepSeek Sparse Attention: a lightweight indexer selects which prior tokens each query attends to, keeping long-context cost sub-quadratic. Pretraining ran 28.5T tokens, up from 23T for the previous generation, with context extended in stages from 32K to 128K to 200K. RL ran on Z.ai's asynchronous "slime" infrastructure. Z.ai reported 77.8 on SWE-bench Verified, 55.1 on SWE-bench Pro, and 56.2 on Terminal-Bench 2.0. Priced at $1.00/$3.20. Technical report: arXiv 2602.15763.

GLM-5.1 (April 7, 2026): long-horizon post-training

Same weights count, same config, new post-training. Z.ai built the release around duration: the model sustains a single task across hundreds of rounds and thousands of tool calls, up to 8 hours in one run. SWE-bench Pro moved 55.1 to 58.4 (the top score at launch, ahead of GPT-5.4 and Claude Opus 4.6), Terminal-Bench 2.0 56.2 to 63.5, CyberGym 48.3 to 68.7. Price rose to $1.40/$4.40, where it has stayed. Full breakdown: GLM-5.1.

GLM-5.2 (June 13, 2026): 1M context

The architectural release. IndexShare reuses one sparse-attention indexer across every four layers, which Z.ai reports cuts per-token FLOPs 2.9x at 1M tokens; that is what made the 200K to 1M jump economical. Multi-token-prediction speculative decoding gained KVShare (reusing the target's KV cache) and rejection sampling, for up to 20% longer draft acceptance. Effort levels arrived, with max as the default. SWE-bench Pro reached 62.1, Terminal-Bench 2.1 reached 81.0, and the Intelligence Index went from 27 to 39. Z.ai also disclosed more reward-hacking behavior than GLM-5.1 and added an anti-hack module. Full breakdown: GLM-5.2 and the GLM-5.2 API guide.

GLM-5.3 (August 14, 2026): scaled post-training, cyber

Same 753B base as GLM-5.2, no new pretraining, a much larger post-training program. Terminal-Bench 3.0 went 4.6 to 28.3, DeepSWE v1.1 46.2 to 66.9, CyberGym 77.2% to 84.5%, ExploitBench 24.4% to 54.4%. The Intelligence Index went 39 to 45, taking the open-weights lead from Kimi K3 at 44. Z.ai says it used the model to find 2,436 vulnerabilities across 269 open-source projects and stood up a disclosure ledger at cvd.z.ai. Weights were held back for safety hardening and landed on Hugging Face on August 25 under the custom GLM-5.3 License. Full breakdown: GLM-5.3.

The license change drew an immediate open request on Z.ai's repo to ship GLM-5.3 under a free-software license like its siblings: "GLM-5.2 and GLM-5.3-Flash use a FLOSS license. Now GLM-5.3 is proprietary. Why the change?" It is still open and unanswered, and the one comment on it links parallel requests filed against Kimi K3, Qwen 3.8, and MiniMax M3, calling it an "unsettling trend." zai-org/GLM-5 #135. For most users the clause is inert, since it binds model-as-a-service operators above $10B in revenue. For anyone whose policy requires an OSI-approved license, GLM-5.2 is the last large model in the family that clears it.

GLM-5.3-Flash (August 26, 2026): the small sibling

A different network: 320B total, 18B active, with vision input the large models lack, 1M context, MIT license. Artificial Analysis scores it 42 on Intelligence Index v4.3, three points under GLM-5.3 and level with Claude Opus 4.8, at about a tenth of GLM-5.3's price. Morph serves it as morph-glm53flash. Full breakdown: GLM-5.3-Flash.

Benchmark progression across the family (Z.ai-reported unless marked)
GLM-5GLM-5.1GLM-5.2GLM-5.3
SWE-bench Pro55.158.462.1not published
Terminal-Bench 2.0 / 2.156.2 (2.0)63.5 (2.0)81.0 (2.1)88.2 (2.1)
Terminal-Bench 3.0n/an/a4.628.3
CyberGym48.368.777.284.5
Intelligence Index v4.3 (AA, independent)28273945
Context200K200K1M1M
Z.ai list, in / out per M$1.00 / $3.20$1.40 / $4.40$1.40 / $4.40$1.40 / $4.40

Terminal-Bench versions are not comparable to each other: 3.0 is a harder revision than 2.1, which is why GLM-5.2 shows 81.0 on 2.1 and 4.6 on 3.0. GLM-5 and GLM-5.1 report Terminal-Bench 2.0; Z.ai's GLM-5.2 docs list GLM-5.1 at 62.0 on 2.1.

GLM 5 vs Claude Opus

The comparison people run is GLM-5.3 against Claude Opus 4.8, because the two land on the same tier and differ 5x on output price. Opus 5 replaced Opus 4.8 at the same $5/$25 in Anthropic's lineup and scores higher, so the honest frame is: GLM-5.3 beats the Opus that was current when it shipped and trails the one that is current now. The earlier rounds of this fight are on GLM-5.2 vs Claude Opus and GLM-5.3 vs Claude.

Independent scores and list prices (Artificial Analysis Intelligence Index v4.3 and Anthropic, checked September 7, 2026)
ModelIntelligence Index v4.3Input / output per MEval output tokensOutput tok/sWeights
GLM-5.345$1.40 / $4.40 (Z.ai)210M73.1 (Z.ai API)Open, GLM-5.3 License
GLM-5.3-Flash42$0.15 / $0.50180M57.6Open, MIT
GLM-5.239$1.40 / $4.40not listed62.3Open, MIT
Claude Opus 4.842$5.00 / $25.00170M58.6Closed
Claude Opus 551$5.00 / $25.00140M53.2Closed
Claude Fable 550$10.00 / $50.00130M61.7Closed

Eval output tokens is the total Artificial Analysis's harness consumed running its Intelligence Index suite; higher means a more verbose model. AA lists Opus 4.8 as deprecated in favor of Opus 5.

Z.ai's benchmark table, GLM-5.3 vs Opus 4.8

Vendor-reported, from the GLM-5.3 model card. GLM-5.3 leads 12 of 16 rows; the three it loses are repo-scale generation (NL2Repo), the long-horizon SWE-Marathon, and Toolathlon.

GLM-5.3 vs Claude Opus 4.8 vs Claude Fable 5 (Z.ai model card, August 2026)
BenchmarkGLM-5.3Claude Opus 4.8Claude Fable 5
Terminal-Bench 2.188.285.088.0
Terminal-Bench 3.028.321.133.7
DeepSWE v1.166.958.069.7
FrontierSWE78.166.588.2
NL2Repo58.069.7not listed
SWE-Marathon v1.142.548.833.1
CyberGym84.578.183.8
ExploitBench54.440.078.0
Toolathlon Verified73.076.274.7
HLE with tools62.557.963.9
Agents' Last Exam (CLI)28.525.723.8
GDPval-AA v21,7691,5881,743

Cost per task, not cost per token

The sticker gap is 3.6x on input and 5.7x on output at Z.ai list, or 4x and 5.7x on Morph. Verbosity eats part of it. GLM-5.2 burned about 43,000 output tokens per Intelligence Index task at max effort, 37,000 of them reasoning, against roughly 16,000 for GPT-5.5. GLM-5.3 kept the habit: 210M output tokens across the AA suite against Opus 4.8's 170M and Fable 5's 130M, and AA measures $2.01 per Index task for it (checked September 7, 2026).

Z.ai's own harness tells the same story from the other side. On its internal Code Bench, GLM-5.3 at high effort scores 31.4% using about 50K output tokens per task, while Opus 4.8 scores 29.5% using 120K. Take those token counts at list price and the output bill per task is $0.22 on GLM-5.3 against $3.00 on Opus 4.8. Take AA's default-effort measurement instead and GLM-5.3 spends 1.4x the tokens Opus 4.8 does, so its per-task cost advantage shrinks from 5.7x to roughly 4x. Either way it is cheaper; the size of the gap depends on the effort setting.

31.4% at ~50K vs 29.5% at 120K
GLM-5.3 (high effort) against Claude Opus 4.8 on Z.ai's internal Code Bench: a higher score at less than half the output tokens. Vendor-internal and unreplicated, but it is the number Z.ai built the launch around.
Z.ai GLM-5.3 docs, August 2026

The one third-party head-to-head

Almost every GLM 5 benchmark in circulation is Z.ai's or Artificial Analysis's. Semgrep ran its own, on a task neither vendor optimizes for: finding insecure direct object reference bugs, using the same dataset and prompt across models. GLM-5.2 driven by a bare prompt in Pydantic AI scored 39% F1, slightly ahead of Claude Code on Opus 4.6 at 37% and well ahead of Claude Code on Opus 4.8 at 28%, at roughly $0.17 per vulnerability found, about a sixth of frontier pricing. Semgrep's own purpose-built multimodal harness beat all of them at 61% with GPT 5.5, which is the honest caveat: harness beats model here. So is theirs: "This is one task, one dataset, one run." Still, it is an independent result on real code, and it lands the same way Z.ai's cyber numbers do. Semgrep, June 2026.

There is a behavioral reason GLM 5 does well on security work specifically, and it is not only capability. Users report it "happily complies with running exploits, reverse engineering and decompiling" where Claude refuses the same request. If your workload is offensive security research, refusal rate is part of the benchmark whether or not anyone publishes it.

Where Claude still wins outright

Vision: every GLM 5 large model is text-only; only GLM-5.3-Flash accepts images. Verified scores: Claude's replicate across harnesses, GLM-5.3's are mostly Z.ai's with one independent datapoint. Peak capability: Opus 5 (63) and Fable 5 (62) sit above GLM-5.3 (60) on the Index, and Fable 5 leads the Z.ai table on ExploitBench (78.0 vs 54.4) and FrontierSWE (88.2 vs 78.1). Effort control cuts the other way: GLM-5.3 exposes low, high, and max levels; cap at high for routine work and the verbosity tax mostly disappears.

Run GLM 5 Locally

Start with the thing that wastes the most time: the Ollama library entry for glm-5 is empty. The page still ranks and still shows 2.3M downloads, but it reads "glm-5 was retired on July 15, 2026" and "No models have been pushed." There is no ollama pull glm-5. Local GLM 5 today means llama.cpp with an Unsloth GGUF, or vLLM/SGLang on datacenter GPUs.

How much RAM and VRAM you need

llama.cpp counts unified memory, system RAM, and VRAM together, so the number that matters is total memory above the file size with headroom for the KV cache. These are Unsloth's published quants for GLM-5.3, with the top-1 accuracy they measure against the full-precision model.

GLM-5.3 GGUF quants (Unsloth, September 2026)
QuantFile sizeTop-1 accuracyPerplexityMemory needed
UD-IQ1_S (1-bit)216.7 GB72.56%4.6130223 GB
UD-IQ2_M (2-bit)238.6 GB78.53%3.7433245 GB
UD-Q2_K_XL253.9 GB80.93%3.5048245 GB
UD-IQ3_XXS (3-bit)281.7 GB84.15%3.2482290-360 GB
UD-Q3_K_XL343.0 GB88.86%2.9107290-360 GB
UD-IQ4_XS (4-bit)365.3 GB90.59%2.8460372-475 GB
UD-Q4_K_XL467.3 GB94.29%2.7006372-475 GB
UD-Q6_K_XL (6-bit)684.4 GB96.59%2.6771570 GB
8-bitnot publishedn/an/a810 GB

GLM-5.3-Flash is the practical option for a workstation: UD-Q2_K_XL is 108.72 GB and wants 115 GB, UD-Q4_K_XL is 199.71 GB and wants 162-210 GB, and BF16 needs 650 GB. For reference, BF16 GLM-5 needs roughly 1,490 GB, which is why nobody runs the large model unquantized on one box.

Do not go below 4-bit if you can avoid it

Unsloth's own table shows the cost: 1-bit keeps 72.56% top-1 accuracy, 2-bit 78.53%, 4-bit 94.29%. Practitioners running the small models report the same cliff from the other side. One Hacker News commenter puts it as "anything less than a 4-bit quant has a tendency to go off the rails," and another running GLM-5.3-Flash at UD-Q2_K_XL reports it "tends toward long thinking loops even for simple tasks," with lower quants lengthening them. Cheaper quants cost you tokens back in reasoning.

Speed you should expect

Unsloth measures GLM-5.3-Flash at UD-IQ3_S on a single B200 at 48.99 tok/s generating at 65,536 tokens of context, and 86.5 tok/s at a 4,096-token prompt with multi-token prediction at n=2. Their note on MTP is worth keeping: the speedup plateaus around 2 draft tokens and more drafts slow it down. For the large model on consumer-class hardware, the numbers people report are lower. One estimate puts a $15,000 512 GB Mac at about 30 tok/s on a GLM-5.3-class model, which is roughly 58M output tokens a month if it never stops.

3 years to break even
A $6,000 local rig at 14 tok/s produces at most 38.5M tokens a month, which is under $164 of GLM-5.3 tokens on the inference market. Buy hardware for privacy, control, or permanence. The token math alone does not get you there.
Hacker News, GLM-5.3 open-weight thread, August 2026

The flags that matter

Sampling: temperature 1.0, top_p 0.95, min_p 0.01 for most work; top_p 1.0 for long agentic runs. Reasoning is always on and cannot be turned off, so the only lever is effort, which llama.cpp passes through the chat template rather than a normal flag:

./llama.cpp/llama-cli \
    --model unsloth/GLM-5.3-GGUF/UD-IQ2_M/GLM-5.3-UD-IQ2_M-00001-of-00006.gguf \
    --temp 1.0 --top-p 0.95 --min-p 0.01 \
    --cache-type-k q4_1 --cache-type-v q4_1 \
    --jinja \
    --chat-template-kwargs '{"reasoning_effort":"low","clear_thinking":true}'

reasoning_effort takes low, high, or anything else, which the template reads as max. clear_thinking defaults to false and Unsloth recommends true for multi-turn chat, which drops old scratchpads instead of re-sending them. Unsloth also had to patch the shipped chat template: GLM writes list indexing as .{id}., which many engines cannot parse, so their GGUFs use [id] instead.

On GPUs

The official vLLM recipe for GLM-5.3 is 8x H200 or 8x H20 at 141 GB each for the FP8 checkpoint, and 8x B200 at 180 GB each if you want the full 1M context. It names three checkpoints, zai-org/GLM-5.3 (FP8), zai-org/GLM-5.3-BF16, and Inferact/GLM-5.3-NVFP4 for Blackwell, and recommends tensor parallel 8, an FP8 KV cache, MTP with 5 speculative tokens, --tool-call-parser glm47, and --reasoning-parser glm45. The recipe warns that synthetic throughput benchmarks under-report real speed because MTP acceptance is low on synthetic traffic.

Serving GLM 5: Throughput and Caching

Open weights mean the same model runs on many hosts, and the hosts are not interchangeable. Artificial Analysis measures GLM-5.3 across providers with a 6.7x spread in output speed and a 2.5x spread in blended price for identical weights.

Capacity is the other axis, and the first-party endpoint is not exempt. Days after the GLM-5.2 launch a paying user filed a capacity report against Z.ai's API with logs: 285 HTTP 429 responses in one day at roughly a 50% failure rate, followed by a second day that worsened through the morning to every message failing in the 11:00 window. zai-org/GLM-5 #83. Launch weeks are the wrong time to depend on a single host for an open-weights model, which is the practical case for keeping a second provider configured.

GLM-5.3 by provider (Artificial Analysis, 10K-token input workload, September 1, 2026)
ProviderOutput tok/sTime to first tokenInput / output per MContext
Databricks299.40.89snot listed ($0.68 blended)1M
DeepInfra761.54s$1.20 / $4.001.05M
Z.ai721.79s$1.40 / $4.401M
Novita711.89snot listed ($1.70 blended)1M
Modular44.89.85s$1.40 / $4.40164K
Morph (morph-glm53-744b)80 on private deploymentsnot measured by AA$1.00 / $3.411M

Morph's 80 tok/s is from the GLM-5.3 744B row of the private-deployment chart on Morph Models, measured with speculators trained on coding traffic; public endpoints run on shared capacity. Artificial Analysis does not measure Morph. GLM-5.3-Flash measures 57.6 tok/s on Z.ai's API.

Why the same weights run at different speeds

A 753B MoE at batch 1 is memory-bound: each token reads every active expert's weights once and does almost no arithmetic per byte, so tokens per second is bandwidth divided by bytes touched. The levers are the ones speculative decoding and quantization pull. GLM-5.2 and GLM-5.3 ship a multi-token-prediction draft head with KVShare, and Z.ai reports up to 20% longer acceptance from its rejection-sampling change. A draft trained on the target's own coding output accepts more tokens per step than a generic one, which is why Morph trains one speculator per model on coding traces. On the memory side, Morph serves GLM-5.3 with NVFP4 expert weights and an FP8 KV cache at tensor-parallel 8; the NVFP4 checkpoint is about 465 GB against roughly 1.5 TB for BF16.

Prefix caching

Every open model on Morph has automatic prefix caching with no cache-write surcharge. Cached input bills at $0.20/M on GLM-5.3 (about 80% off) and $0.02/M on GLM-5.3-Flash (about 85% off). Agent loops get this for free: each turn re-sends the previous turns verbatim, so everything but the newest turn hits. Two controls: prompt_cache_key pins a conversation to the worker holding its prefix, and cache_ttl sets retention from 5 minutes to 24 hours. Details in the caching docs.

The cached rate is the one to budget against, not the output rate. A practitioner working the numbers on Hacker News argues that for agentic coding roughly 90% of the bill is cached input, and that it grows with the square of session length, because every turn re-sends every prior turn. Their worked case: sessions running out to 1M context can push past 1B cached input tokens in a day, which at the $0.26 per million cached rate is $260 a day from cache reads alone. That is also the argument for clear_thinking and for capping effort. Shorter sessions and dropped scratchpads shrink the quadratic term, not just the output line.

GLM 5 on Morph, per 1M tokens (live billing rates)
ModelInputCached inputOutputStandby / Batch (50% off)
GLM-5.3 (morph-glm53-744b)$1.00$0.20$3.41$0.50 / $0.10 / $1.705
GLM-5.3-Flash (morph-glm53flash)$0.10$0.02$0.35$0.05 / $0.01 / $0.175

Standby: send service_tier: "standby" and pay half; the request runs on spare capacity and sheds a fast 429 when a region is busy, billing nothing on rejection. The Batch API bills at the same rate with a 24-hour window. See the standby docs.

Serving Gotchas People Actually Hit

These are open or recently closed issues in the two engines most people serve GLM 5 on. They are the difference between a working deployment and a model that looks broken.

enable_thinking does nothing on GLM-5.3, and breaks the output

GLM-4.5 and GLM-4.6 honored chat_template_kwargs: {"enable_thinking": false}. GLM-5.3's template never reads that name. Its only knobs are reasoning_effort and clear_thinking. vLLM's reasoning parser still gates extraction on the old kwarg, so a client that sends it gets the worst of both: the model thinks anyway, and the parser stops separating the scratchpad, which lands in message.content with a dangling </think> in front of the answer. A vLLM maintainer's reply on the issue is blunt: "GLM-5.3 officially no longer supports enable_thinking." If you are porting a GLM-4.x client, delete that kwarg first. vLLM #54744.

MTP speculative decoding can silently destroy quality

Multi-token prediction is the headline speedup on GLM-5.2 and GLM-5.3, and it is configuration-sensitive. One report has GLM-5.2-FP8 on 16 H200s with expert parallelism scoring 95.68% on GSM8K without MTP and 1.06% with a single speculative token, which is garbage output rather than a small regression. A vLLM committer could not reproduce it on a single-node 8x B300 deployment, where the same config passes at 94.0%; the reporter's setup needed a patch for dual-batch overlap. The lesson is not that MTP is broken, it is that MTP acceptance and correctness are worth an eval gate on your exact topology before you trust the throughput number. vLLM #46834.

AMD paths lag the NVIDIA ones

On 8x MI355X with the Quark MXFP4 checkpoint, GLM-5.2 MTP starts cleanly and then accepts zero draft tokens, so speculative decoding is pure overhead while it drafts at 40-230 tok/s; disabling expert parallelism in the same config hits an illegal memory access. vLLM #52833. GLM-5.3's Quark MXFP4 checkpoint has a separate loading failure on ROCm nightlies (vLLM #54900). SGLang is landing tuned MI355X and gfx950 recipes for GLM-5.3-Flash, so this is moving, but check the date on any AMD recipe you copy.

GLM-5.3-Flash is a different architecture, not a smaller GLM-5.3

Flash uses gated linear attention layers that vLLM reports as Glm5NextTextLinearAttention, with convolution and forget-gate weights the standard MoE loader does not know about. People following the B200 recipe on the nightly wheel hit a hard load failure on self_attn.k_conv1d; a maintainer's answer was to use the docker image from the recipes site instead, because "nightly doesn't have glm 5.3 flash now." vLLM #54062. Treat Flash as its own port, not a config change.

It sometimes answers in Chinese, and has since GLM-5.1

The longest-running user complaint on Z.ai's own tracker is spontaneous language switching: mid-response, in an all-English context with an English system prompt, the model emits a Chinese sentence and then resumes English. It was filed against GLM-5.1 in April 2026, confirmed by several users on GLM-5.2, and filed again against GLM-5.3-Flash in August, where the reporter notes it survives a context reset. A Z.ai maintainer's reply asks whether quantization is the cause and says the rate on GLM-5.2 "should be very low." One report narrows the mechanism usefully: the reasoning trace stayed in English while only the final output switched, and the model denied switching when asked. zai-org/GLM-5 #54 and #142. If you ship user-facing output, validate the language.

Tool-call JSON breaks on some providers

An open report has GLM-5 through NVIDIA NIM returning tool-call arguments with truncated JSON, for example a query string with no closing brace, sending OpenCode into parse-and-retry loops across several different tools in one session. A commenter on the thread raises the obvious suspect: "Are you running it with MTP? I remember having a lot of tool calling issues when using MTP." zai-org/GLM-5 #15. Same weights, same prompt, different serving stack, invalid output.

Sparse attention makes cache offload harder than it looks

SGLang shipped a fix for GLM-5.3-Flash returning corrupted output after HiCache moved a cached prompt to host memory and back: KV values and recurrent state were restored, but the sparse-attention index buffers that choose which tokens to attend to were not, so the model resumed with the wrong attention state even with speculative decoding off. A second defect let two divergent prompt suffixes share a compressed index row. SGLang #38212. If you are building your own KV offload tier for a DSA model, index buffers are part of the state.

Why the same model feels different per provider

The list above is the mechanism behind a complaint that shows up constantly in open-weights threads. As one practitioner puts it, "for open-weight models the provider's setup impacts performance so you can have different experiences with the same model at the same quantization from different providers," including a middleware update introducing a defect that degrades quality. Benchmarks measure a checkpoint. You are buying a deployment.

GLM 5 API: Morph and Z.ai

Both endpoints are OpenAI-compatible. Switching hosts is a base URL and model name change. On Morph the model ids are morph-glm53-744b for GLM-5.3 and morph-glm53flash for GLM-5.3-Flash; one key works for every model in the lineup.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.morphllm.com/v1",
    api_key="YOUR_MORPH_API_KEY",
)

resp = client.chat.completions.create(
    model="morph-glm53-744b",  # or "morph-glm53flash"
    messages=[
        {"role": "user", "content": "Find the race condition in this queue and fix it."},
    ],
)
print(resp.choices[0].message.content)

Claude Code

Morph serves the Anthropic Messages API at /v1/messages for every open model, so Claude Code runs on GLM-5.3 with env vars and no proxy. ANTHROPIC_MODEL remaps sonnet and opus; ANTHROPIC_SMALL_FAST_MODEL remaps the haiku background calls.

export ANTHROPIC_BASE_URL="https://api.morphllm.com"
export ANTHROPIC_AUTH_TOKEN="YOUR_MORPH_API_KEY"
export ANTHROPIC_MODEL="morph-glm53-744b"
export ANTHROPIC_SMALL_FAST_MODEL="morph-glm53flash"
claude

Use /effort high rather than max for routine work to keep the reasoning-token bill down. Full setup, including the settings.json form, in the Claude Code guide.

Cline, OpenCode, Cursor

Set the OpenAI-compatible provider base URL to https://api.morphllm.com/v1, the model to morph-glm53-744b, and the context window to 1000000. Leave image support unchecked for GLM-5.3; enable it for morph-glm53flash, which accepts images.

Z.ai first-party

Base URL https://api.z.ai/api/paas/v4, model glm-5.3 (or glm-5.3-flash), at $1.40/$4.40 per million tokens. Z.ai's Anthropic-format endpoint for Claude Code is https://api.z.ai/api/anthropic. Older ids glm-5.2, glm-5.1, and glm-5 remain callable on Z.ai; on Morph the older morph-glm52-744b id resolves to the GLM-5.3 stack.

The subscription route

Z.ai also sells the GLM Coding Plan, a flat monthly subscription rather than per-token billing, starting at $18 a month. It is metered in credits on two windows: Lite gets 2,000 credits per 5 hours and 10,000 a week, Pro 12,000 and 60,000, Max 28,000 and 140,000. All tiers cover GLM-5.3 and GLM-5.3-Flash, and requests for older ids route to the newer versions automatically. The plan is restricted to supported tools, which include Claude Code, Cline, and OpenCode; Claude Code and Goose point at https://api.z.ai/api/anthropic and the rest use https://api.z.ai/api/coding/paas/v4. Subscriptions are the cheaper path for one person coding all day; per-token APIs are the cheaper path for anything programmatic, batch, or bursty, and they are the only path with a standby or batch discount.

GLM 5: Pros and Cons

Strengths
  • Top open-weights tier: GLM-5.3 scores 45 on the independent Intelligence Index v4.3, #1 of 112 open-weights models, ahead of Kimi K3 (44) and Claude Opus 4.8 (42), checked September 7, 2026
  • Price held at $1.40/$4.40 list across three releases while capability rose from Index 28 to 45
  • Open weights for every version: MIT for GLM-5 through GLM-5.2 and GLM-5.3-Flash, the GLM-5.3 License for GLM-5.3
  • 1M context since GLM-5.2 via IndexShare (2.9x per-token FLOP cut at 1M)
  • Runs natively in Claude Code through the Anthropic Messages API on Morph or Z.ai
  • Effort levels (low/high/max) give a real lever on reasoning spend
  • GLM-5.3-Flash covers vision and high-volume work at about a tenth of the large model's price
Limitations
  • Verbose: 210M eval output tokens for GLM-5.3 vs a 120M open-weights median; max effort is the default
  • Nearly every launch benchmark is Z.ai-reported; the independent record is Artificial Analysis plus Semgrep's single IDOR run
  • Reasoning cannot be disabled, and enable_thinking: false leaks the scratchpad into content on vLLM (issue 54744)
  • MTP speculative decoding is topology-sensitive: one report shows GSM8K falling from 95.68% to 1.06% on a 16-GPU GLM-5.2-FP8 setup
  • No Ollama build: the glm-5 library entry was retired on July 15, 2026 with no models pushed
  • Spontaneous Chinese output in English contexts, reported against GLM-5.1, GLM-5.2, and GLM-5.3-Flash and still open
  • Large models are text-only; only GLM-5.3-Flash takes images
  • GLM-5.3 weights carry a custom license with a $10B-revenue security-review clause, not MIT
  • Self-hosting the 753B model is an 8-GPU job even at NVFP4 (~465 GB)
  • GLM-5.2 disclosed more reward-hacking behavior than GLM-5.1; eval-gate agents with filesystem or test access
  • Speed varies 6.7x across hosts for the same weights, so a bad host makes the model look slow

FAQ

What is GLM 5?

Z.ai's family of open-weight coding models: GLM-5 (February 11, 2026), GLM-5.1 (April 7), GLM-5.2 (June 13), GLM-5.3 (August 14), and GLM-5.3-Flash (August 26). The large models are 753B MoEs with about 40B active per token; Flash is 320B/18B with vision. Also written GLM5 or GLM-5.

Is GLM 5 open source?

Open weights. GLM-5, 5.1, 5.2, and 5.3-Flash are MIT on Hugging Face. GLM-5.3 is public since August 25 under the custom GLM-5.3 License, which is MIT-like plus a security-review requirement for model-as-a-service operators above $10B in revenue. Training data and code are not released. The change drew an open request on Z.ai's repo to put GLM-5.3 back under a free-software license, still unanswered. If your policy needs an OSI-approved license on the large model, GLM-5.2 is the last one that qualifies.

Is GLM 5 better than Claude?

GLM-5.3 scores 45 on the independent Intelligence Index v4.3 against Opus 4.8's 42, Opus 5's 51, and Fable 5's 50 (checked September 7, 2026), and leads Opus 4.8 on 12 of 16 rows of Z.ai's own table. So: above the Opus that was current at launch, below the current Anthropic top tier, at a fifth of Opus's output price and with no vision. Details on GLM-5.3 vs Claude.

How much does GLM 5 cost?

Z.ai list: $1.40/$4.40 per million tokens for GLM-5.1 through GLM-5.3, $1.00/$3.20 for the original GLM-5, $0.15/$0.50 for Flash. Morph: $1.00/$3.41 for GLM-5.3 and $0.10/$0.35 for GLM-5.3-Flash, half that on standby or batch.

How do I call the GLM 5 API?

OpenAI-compatible on both hosts. Morph: https://api.morphllm.com/v1 with morph-glm53-744b or morph-glm53flash. Z.ai: https://api.z.ai/api/paas/v4 with glm-5.3. Claude Code works against either through the Anthropic Messages API.

Who owns GLM 5? Which company makes it?

Z.ai, the Beijing lab previously known as Zhipu AI and one of the group Chinese press calls the "six AI tigers." The GLM line starts with a March 2021 paper on General Language Model pretraining and reaches consumers with ChatGLM in March 2023. Weights ship under the zai-org org on Hugging Face. Nobody else owns the model, but anyone can host the open weights, which is why the same GLM 5 runs on Z.ai, Morph, DeepInfra, Novita, and Databricks.

When did GLM 5 come out?

February 11, 2026. Then GLM-5.1 on April 7, GLM-5.2 on June 13, GLM-5.3 on August 14 with the weights following on August 25, and GLM-5.3-Flash on August 26.

How much VRAM and RAM does GLM 5 need?

For GLM-5.3 under llama.cpp, count RAM and VRAM together: about 223 GB at 1-bit, 245 GB at 2-bit, 372-475 GB at 4-bit, 810 GB at 8-bit. GLM-5.3-Flash needs about 115 GB at 2-bit and 162-210 GB at 4-bit. On GPUs, the official vLLM recipe is 8x H200 or H20 at 141 GB each, or 8x B200 at 180 GB each for the full 1M context. Full details in Run GLM 5 Locally.

Can I run GLM 5 with Ollama?

No. The glm-5 entry in the Ollama library was retired on July 15, 2026 and carries no pushed models, despite still showing 2.3M downloads. Use llama.cpp with an Unsloth GGUF, or vLLM or SGLang on GPUs.

How do I turn off thinking on GLM 5?

You cannot, and trying breaks the response. GLM-5.3's template never reads enable_thinking, so sending it makes the model think anyway while vLLM's parser stops extracting, dumping the scratchpad into message.content. Use reasoning_effort set to low or high, and clear_thinking to drop old scratchpads between turns.

What is GLM 5's context window?

200K tokens on GLM-5 and GLM-5.1, 1M on GLM-5.2, GLM-5.3, and GLM-5.3-Flash, with 128K max output. Morph serves the full 1M on both models it hosts.

Related Articles

Operator answer

The balanced middle

Independent results place it between Flash efficiency and Kimi scale, without forcing every task onto either extreme.

Best fits

  • Coding and cyber work
  • Capability per dollar
  • Faster output than very large frontier MoE models

Escalate or test carefully

  • It is not the cheapest option
  • It is not the highest capability option
  • Always on reasoning can increase token use
Serving architecture

Cache the full agent session

Long coding sessions reuse system prompts, repository context, tool output, and prior turns. A useful production stack tiers that cache across GPU memory, CPU memory, and NVMe. GPU only cache sizing misses much of the cost per task opportunity.

Hardware guidance

Start with GB300 NVL72

Choose hardware around required speed per active user, then measure total capacity inside that latency target. Large batch throughput alone can hide a slow agent experience.

Dedicated inference planner

Plan a GLM-5.3 744B endpoint

Turn your team size and agent workload into a capacity estimate. Then validate the recommendation with your own traces.

Workload economics
Per model decision
900M
tokens per month
100 tok/s
required generation
$797
serverless per month
$66,576
dedicated per month
Sizing review required

An exact Morph capacity measurement is required before recommending a dedicated plan.

GB300 NVL72 is the compatible public platform. Dedicated capacity is invoiced monthly at the beginning of the month. Tokens are not billed separately.

Difference from serverless: $65,779 more per month.

Planning the smaller model instead? The GLM-5.3-Flash planner and B200 sizing are on the GLM-5.3-Flash page. For the general method, see the capacity calculator and the dedicated inference benchmarks.

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.

Talk to us about a private deployment

Run GLM-5.3 on tuned kernels

GLM-5.3 at $1.00/$3.41 and GLM-5.3-Flash at $0.10/$0.35 per million tokens, full 1M context, prefix caching on by default, half price on standby. One OpenAI-compatible key across Kimi K3, GLM-5.3, GLM-5.3-Flash, and DeepSeek V4 Flash.

Sources