~/satyajit

GLM-5.3: the same base model, a month of post-training, and a cyber capability nobody planned

mdjsonmcp

2026-08-14 · 15 min · agents · post-training · reinforcement-learning · security

Why read this

Solidtop 85%

Diffs the 5.2 and 5.3 configs, counts active parameters from all 141 shard headers (41.25B, not 40B) and shows Unsloth's "2-bit" figures mix two files.

  • Original analysis
  • Widely used
  • Concrete numbers to act on

Training & RLNeeds a workstation GPUCustom licencePractitioner model

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
2 of 3: Open code or weights with real limits
Will I understand it?
1 of 3: Partial mechanism
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
1 of 3: Relevant for months
Does it affect many?
2 of 3: A widely used model, tool or lab release
Only here?
2 of 3: A teardown or measurement few others did

Score 57 of 100, ranked 290 of 454 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Z.ai's GLM-5.3 post opens with a sentence most labs would bury:

Scaling post-training is all we did for GLM-5.3.

It is the same base model as GLM-5.2. No new pretraining run, no new architecture, nothing to report about parameter counts because none of them changed. What changed is a month of additional post-training on the stack they had already built — IndexShare for long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — pointed at more environments, more diverse tasks, and more compute.

That makes the release unusually easy to read. Every number below is a post-training delta, which is a rarer thing to be able to say than it sounds.

zai-org/GLM-5.3@aca966e · snapshot 2026-09-22
parameters
753.33B
repo size
755.68 GB
architecture
GlmMoeDsaForCausalLM
task
text-generation
library
transformers
license
other
safetensors
141 shards
largest file
5.37 GB
files
155
downloads
1.0M
likes
1.9K
languages
en, zh
parameters by dtype
BF162.10BF3219.5KF8_E4M3751.23B

repo last modified 2026-09-04

What a month of post-training bought

same base model · post-training only
Terminal Bench 2.181 → 88.2 · 1.09×
Terminal Bench 3.04.6 → 28.3 · 6.15×
6.2x — the largest relative move on the board, off a floor so low that 5.2 was essentially not playing.
DeepSWE (v1.1)46.2 → 66.9 · 1.45×
NL2Repo48.9 → 58 · 1.19×
ProgramBench (Almost Solved)9.5 → 19 · 2.00×
FrontierSWE67.5 → 78.1 · 1.16×
SWE-Marathon (v1.1)19.4 → 42.5 · 2.19×
PostTrainBench31.7 → 39.8 · 1.26×
CyberGym77.2 → 84.5 · 1.09×
ExploitBench24.4 → 54.4 · 2.23×
The 'more than doubles' claim in the post. It checks out — 2.23x.
Toolathlon Verified59.9 → 73 · 1.22×
AutomationBench (v1.0.6)26.2 → 48.2 · 1.84×
Agents' Last Exam (ALE-CLI)23.8 → 28.5 · 1.20×
The smallest relative gain of the sixteen, at 1.20x.
HLE w/ Tools54.7 → 62.5 · 1.14×
GLM-5.2GLM-5.3

Fourteen benchmarks where both models are scored, and GLM-5.3 is ahead on every one of them. That is a cleaner result than it looks, because the two models share a base: whatever moved here was moved by post-training, not by a new pretraining run. The spread is the interesting part — a 1.20× gain on Agents’ Last Exam and a 6.2× gain on Terminal-Bench 3.0 are not the same kind of claim, and the biggest multiples all sit on benchmarks where GLM-5.2 started near the floor.

Fourteen benchmarks where both models are scored, GLM-5.3 ahead on all fourteen. The two that dominate the story:

At the other end, Agents' Last Exam moves 23.8 → 28.5, a 1.20× gain — the smallest of the sixteen. The gains are real and they are not uniform.

Where the environments came from

The part of the post I found most interesting is not a benchmark. Z.ai describes the bottleneck moving off the model entirely:

As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.

Their answer is to synthesize the environments, and for a subset of tasks the reward signal too. Research agents collect task patterns from real work and turn them into runnable long-horizon environments with multi-step dependencies and hidden state; a judge agent then attempts each task to confirm it is actually solvable. Verifiers are synthesized without access to the reference solution, and solver trajectories are used to find and close reward shortcuts. A verifier that passes oracle, no-op and unsolved-state checks produces a binary reward they consider reliable enough to train on directly.

That trio of checks is the load-bearing detail. An oracle check catches a verifier that rejects correct solutions; a no-op check catches one that accepts doing nothing; an unsolved-state check catches one that was already satisfied before the agent started. Those are the three ways a synthesized reward usually turns out to be worthless, and they are checkable without a human reading the task.

Z.ai is direct that this is not yet automatic: the pipelines "still require a meaningful amount of human-in-the-loop work."

The environments themselves are aimed at something closer to a job than an exercise. Their example is an ML infrastructure task where the model gets the same working environment as an engineer — compute clusters, storage, internal documentation, codebases, experiment results — and has to diagnose bottlenecks across a training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup without breaking correctness. Some tasks, they say, represent several days of work for an experienced engineer.

The honest reading of the table

16 benchmarks · 8 models
GLM-5.2
16
16W–0L
Kimi K3
10
4
10W–4L
DeepSeek-V4 Pro
7
2
7W–2L
Qwen3.8-Max
12
12W–0L
Opus 4.8
13
3
13W–3L
Fable 5
6
9
6W–9L
GPT-5.6 Sol
4
9
4W–9L
open weightsclosedrows GLM-5.3 loses

Z.ai calls GLM-5.3 “the most capable open-weights model for coding,” and against the open field the table supports it: 16–0 over its own predecessor, 12–0 over Qwen3.8-Max, 10–4 over Kimi K3, 7–2 over DeepSeek-V4 Pro. The closed frontier is a different story — 6–9 against Fable 5 and 4–9 against GPT-5.6 Sol. Counted across the whole field rather than pairwise, GLM-5.3 has the top score on 3 of the 16 rows: CyberGym, AutomationBench and GDPval-AA v2. Both readings are in the same table, and the phrase “open-weights” is doing the work that makes the headline true.

"The most capable open-weights model for coding" is a carefully worded claim and it survives checking. Against open models GLM-5.3 is 16–0 over GLM-5.2, 12–0 over Qwen3.8-Max, 10–4 over Kimi K3, and 7–2 over DeepSeek-V4 Pro. Against the closed frontier it is 6–9 versus Fable 5 and 4–9 versus GPT-5.6 Sol.

Counted across the whole field rather than pairwise, GLM-5.3 holds the top score on 3 of 16 rows — CyberGym, AutomationBench, and GDPval-AA v2. Publishing a table where you lead three rows and lose thirteen is not the normal shape of a launch post, and the word doing the work in the headline is open-weights.

Two smaller notes on the table. Thirteen of the 112 non-GLM-5.3 cells are blank, so several head-to-head records rest on fewer than sixteen rows — the DeepSeek-V4 Pro comparison is nine rows, not sixteen. And ExploitGym is reported as 2h / 6h pairs rather than a single score, so any ranking of that row depends on which budget you pick.

Token efficiency is the better result

The claim I'd have led with is the one about cost, not capability.

Scatter plot of accuracy against average output tokens per task on Z.ai Code Bench, with four models each plotted at several effort levels. GLM-5.3 rises steeply from about 24.7 percent at 48K tokens to 34.5 percent at 75K. GLM-5.2 sits lower and further right. Claude Fable 5 is highest overall, reaching 39.5 percent at about 115K tokens. Claude Opus 4.8 reaches 29.5 percent at about 120K tokens.
Accuracy against output tokens on Z.ai Code Bench v1.0, run inside Claude Code 2.1.207. Up and to the left is better. (z.ai, GLM-5.3 launch post.)

At Max effort GLM-5.3 reaches 34.5% at roughly 75K output tokens per task, against GLM-5.2's 23.4% at 96K — more accurate and cheaper, which is the direction that rarely happens on its own. At High effort it reaches 31.4% at around 50K tokens.

Z.ai compares that last figure to "Claude Opus 4.8 at 29.5% with 120K," which is true and worth reading precisely: 29.5% is Opus's Max effort, not its High. The comparison is GLM-5.3's High against Opus's Max. As an efficiency-frontier argument that is legitimate — the whole point of the chart is that the curves sit in different places — but it is not a like-for-like row.

The post is also straightforward that the frontier still belongs to someone else: "GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort."

One detail in the figure's own subtitle deserves attention: the benchmark was evaluated on Claude Code 2.1.207. Every model in that chart was scored through a competitor's harness. Given how much a harness shapes agent results, holding it fixed across models is the right call, and it is unusual to see it stated on the chart itself.

The cyber result, and what it actually says

This is the part Z.ai describes as a surprise:

As we scaled post-training, cyber capability developed faster than we expected.

They added vulnerability discovery data and environments to the training mix expecting the model to get better at finding and reasoning about flaws. What they report instead is that it began reasoning across multiple stages of exploitation and forming coherent plans for complete chains.

up the exploitation chain1 of 3
CyberGymVulnerability discovery: locate a real defect in a real codebase.
GLM-5.3
84.5%
GLM-5.2
77.2%
Fable 5
83.8%
GPT-5.6 Sol
83.6%
1.09× over GLM-5.2 · -1% behind GPT-5.6 Sol

GLM-5.3 leads the whole field here, and this is the row the SOTA claim rests on. The margin is thin — 84.5 against 83.8 and 83.6 — but it is a lead, and it is over closed models.

Both of Z.ai’s cyber claims are accurate, and they point in different directions. GLM-5.3 really is top of the field at finding vulnerabilities, and its gains really are largest further up the chain when measured against GLM-5.2. But the gap to the closed models widens at exactly the same rate: level at discovery, twenty-plus points behind at exploitation, less than half the throughput at full chains. The capability that emerged is real; the ranking it earned is confined to the first rung.

Both cyber claims in the post are accurate. GLM-5.3 does hold the top CyberGym score at 84.5 — narrowly, over Fable 5 at 83.8 and GPT-5.6 Sol at 83.6, but it is a genuine lead over closed models. And its gains really are largest further up the chain when measured against GLM-5.2: 2.23× on ExploitBench, 3.3× on ExploitGym at 6h.

Put the two sentences next to each other and they suggest a model leading at exploitation. The same table says otherwise. One rung up from discovery, GLM-5.3 sits at 54.4 on ExploitBench against 78 and 76.5. At full chains under a six-hour budget it clears 130 problems against 247 and 293. The gap to the closed models widens at precisely the rate the capability is described as growing.

Alongside this, Z.ai published a disclosure ledger: 2,436 findings tracked, 53 publicly disclosed, 2,383 still under embargo, 1,097 rated critical or high, across 269 open-source projects. The detail that stops you is the age distribution — the oldest flaw was introduced in 1981, and on average a vulnerability had been sitting in a codebase for 26.6 years before it was found. That is a claim about the state of open-source security as much as about the model.

It is also the context for the release schedule. The weights are not out yet:

We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.

A two-week hold between announcement and weights, explicitly attributed to safety evaluation, is a reasonable response to having just demonstrated automated vulnerability discovery at scale. It also means nobody outside Z.ai can check any of the above yet.

What to take from it

The headline result is not the benchmark table, which shows a strong open-weights model that trails the closed frontier — a familiar position. It is the claim that a month of post-training on a fixed base moved fourteen benchmarks, several of them by more than 2×, with output token counts going down. If that reproduces when the weights land, the interesting variable in this release is the environment synthesis pipeline, not the model.

The caveats: Z.ai Code Bench is private, so its numbers cannot be independently checked by construction — a deliberate anti-contamination trade with a real cost. Thirteen table cells are blank. The weights are two weeks out. And the cyber capability that is described as emergent is, on the evidence published alongside it, still a discovery capability rather than an exploitation one.

The weights land, and confirm it

The two weeks passed. zai-org/GLM-5.3 is on Hugging Face, and the central claim of this article — same base model as GLM-5.2, nothing to report about parameter counts because none of them changed — stops being something to take on Z.ai's word and becomes something to check against a file.

So: fetch both config.jsons and diff them.

config.json · GLM-5.2 vs GLM-5.3 · 56 keys
15 of 54 identical keys shown — the architecturally load-bearing ones
architecturesGlmMoeDsaForCausalLM
model_typeglm_moe_dsa
num_hidden_layers78
first_k_dense_replace3 (3 dense, 75 MoE)
hidden_size6,144
n_routed_experts256
num_experts_per_tok8
n_shared_experts1
moe_intermediate_size2,048
max_position_embeddings1,048,576 (1M)
num_nextn_predict_layers1 (MTP)
vocab_size154,880
topk_method / scoring_funcnoaux_tc / sigmoid
routed_scaling_factor2.5
indexer_types78 entries, full/shared pattern
54 of 56 keys identical2 differ

At launch this article could only take Z.ai’s word for “same base model as GLM-5.2.” Now the two configs can be diffed directly: 56 top-level keys in the union of both files, 54 byte-identical, and exactly 2 different — quantization_config, which 5.3 ships and 5.2’s file simply does not have the key for, and transformers_version, a library-metadata bump with no architectural meaning. Every number that describes the model itself — layer count, expert count, hidden size, context length, the routing scheme, the indexer pattern — did not move. The claim holds at the level of the file, not just the press release.

Fifty-six top-level keys across the union of the two files. Fifty-four are byte-identical. The two that differ are quantization_config — 5.3 ships an fp8 block-quantization spec that 5.2's file doesn't have the key for at all — and transformers_version, a library-metadata bump from 5.12.0 to 5.15.0 with no architectural content. Everything that describes the model itself — 78 layers, 256 routed experts at top-8, a 6,144 hidden size, the 1M-token context, the DeepSeek-V3-style noaux_tc/sigmoid aux-loss-free router — is unchanged.

One architectural detail is worth pulling out on its own: indexer_types, a 78-entry list that alternates in groups of four — three full layers, then full, shared, shared, shared repeating. A full layer computes its own sparse-attention token selection; a shared layer skips that computation and reuses a nearby full layer's index instead, and index_share_for_mtp_iteration: true extends the same reuse into the speculative-decoding layer. Twenty-one of the 78 layers do the full computation; the other 57 borrow it. It is the same cross-layer index-reuse idea Tencent's Hy4-preview ships under the name IndexCache — whether one release directly inspired the other isn't stated in either model card, but the two arrived within about two weeks of each other, which is a reasonable signal that reusing a neighboring layer's sparse-attention index is becoming a standard move for anyone building on DeepSeek Sparse Attention, not a one-off trick specific to either model.

Counting the parameters directly. Z.ai's own materials never state a parameter count for GLM-5.3 — the whole point of the launch was that nothing changed — so the number to check is Unsloth's: their docs describe GLM-5.3 as "a new 744B parameter (40B active) model." That number can be verified without downloading 1.5TB of weights. The repo's model API reports a per-dtype element count — BF16, F8_E4M3, F32 — summing to exactly 753,329,940,480 total parameters, but that total alone can't say how many are active per token, because it doesn't know which tensors are routed experts. For that, the shapes are needed, and the checkpoint ships as 141 safetensors shards. Each shard's header — a tensor-name-to-shape map — sits in the first few kilobytes of the file, readable with two HTTP range requests per shard (8 bytes for the header length, then that many bytes of JSON) without touching the weight data at all. Unioning all 141 headers gives the shape of every one of the model's 118,629 tensors.

Summed directly, every tensor's element count — including the fp8 quantization scale factors, which live in the same shards but aren't real parameters — comes to 753,375,793,584, about 45.85 million over the API's figure. Dropping every weight_scale_inv tensor closes that gap to exactly zero: 753,329,940,480, matching Hugging Face's own count to the last digit. That match is the useful part — it means the shape data is being read correctly, and the active-parameter number built from it can be trusted the same way.

Of those 753.33B parameters, only the mlp.experts.* tensors are conditionally active — 8 of each layer's 256 routed experts fire per token, everything else (attention, the shared expert, embeddings, the router, layernorms) runs in full. Layer index 78 turns out to be a 79th layer beyond the 78-layer backbone: it carries its own eh_proj/enorm/hnorm tensors, the signature of a DeepSeek-V3-style multi-token-prediction module, with its own attention block and its own 256 routed experts. Splitting it out:

totalactive (8/256 of routed experts)
78-layer backbone743,377,019,904 (743.38B)41,250,530,304 (41.25B)
+ MTP layer753,329,940,480 (753.33B)41,841,764,352 (41.84B)

The backbone-only total, 743.38B, lands within 0.1% of Unsloth's stated 744B — good evidence their figure is backbone-only, and good evidence the shape-based method is sound. The active count is the one that doesn't close as cleanly: 41.25B backbone-only, or 41.84B counting the MTP layer's own routing, both 3–5% above the stated "40B active." I'd trust the number computed here over the rounder one — it comes directly from the tensor shapes the model actually ships with, cross-checked exactly against an independent total from Hugging Face's own API, where "40B" reads like a round marketing figure that may predate final shapes or simply round down for the headline.

What else the release confirms. The license is broad — free use, modification, redistribution, commercial deployment, no royalty — with one specific carve-out: an organization running a "Model as a Service" business whose aggregate revenue (with affiliates) exceeds $10 billion over any trailing 12 months must pass a Z.ai security review before commercial use. That threshold excludes essentially everyone who might read this. Deployment support arrived unusually wide for a day-zero release: the model card lists SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth, and Ascend-NPU stacks (vLLM-Ascend, xLLM, SGLang) as supported serving paths at launch, not "coming soon." And reasoning_effort accepts low, high, or max, defaulting to max — but per Unsloth's docs, "thinking cannot be disabled." Unlike some 2026-era releases that ship an explicit no-think switch for latency-sensitive use, GLM-5.3 has no off position; every call reasons, at a budget you choose.

The GGUF quants, and a "2-bit" that means two different files

Unsloth shipped Dynamic GGUF quants the same day, twelve variants from a full BF16 checkpoint down to 1-bit. Their docs page makes two accuracy claims — "Dynamic 1-bit GGUF reaches ~76% top-1 accuracy while being 85% smaller. Dynamic 2-bit reaches ~81% accuracy while being 83% smaller" — and then, in the walkthrough further down the same page, recommends a specific file for people actually trying to run it.

unsloth/GLM-5.3-GGUF · day-zero quants
BF16
1508.0GB−0%
Q8_0
801.4GB−47%
UD-Q6_K_XL
684.4GB−55%
UD-Q5_K_XL
562.5GB−63%
UD-Q4_K_XL
467.3GB−69%
UD-IQ4_XS
365.3GB−76%
UD-Q3_K_XL
343.0GB−77%
UD-IQ3_XXS
281.7GB−81%
UD-Q2_K_XL
253.9GB−83%
UD-IQ2_M
238.6GB−84%
UD-IQ1_M
228.5GB−85%
UD-IQ1_S
216.7GB−86%
UD-Q2_K_XL — 254GB, −83%, ~81% accuracy (Unsloth’s own pairing)UD-IQ2_M — 239GB, −84% (the file Unsloth’s walkthrough actually recommends)

Both highlighted rows are real files Unsloth ships, and both get called “2-bit” on the same docs page. UD-Q2_K_XL is the one the accuracy claim belongs to — 254GB, −83% smaller than BF16, and the “~81% accuracy” figure is stated right next to it. UD-IQ2_M is a different, smaller, importance-quantized file — 239GB, −84% — that the walkthrough later recommends instead, for “the best balance of accessibility and accuracy.” Neither number is wrong. The promotional shorthand that reads “239GB, −83% smaller, ~81% accuracy” as one fact is quoting the size of one file and the shrink-and-accuracy pair of its 15.3GB-larger neighbor.

Sum the repo's own file listing by folder and the ladder is exactly what the docs describe at the ends: BF16 at 1,508.0GB, and the smallest 1-bit variant at 216.7GB, an 86% shrink. The middle is where it gets specific. UD-Q2_K_XL is 253.9GB — a 1 − 253.9 / 1508.0 ≈ 83.2% reduction, which is the file the "~81% accuracy... 83% smaller" sentence is describing. UD-IQ2_M is a different, smaller, importance-quantized file at 238.6GB — a 1 − 238.6 / 1508.0 ≈ 84.2% reduction — and it's this second file the walkthrough actually recommends: "We will be utilizing the 239GB UD-IQ2_M quant for the best balance of accessibility and accuracy." Read the two sentences as one fact — 239GB, 83% smaller, ~81% accuracy — and they don't describe any single file that exists. The 239GB figure and the 83%/81% figures belong to two different quants, 15.3GB apart, both of which Unsloth calls "2-bit" a few paragraphs apart on the same page. Neither number is fabricated; the shorthand just quietly stitches one file's size to its neighbor's accuracy claim.

The hardware side holds up better than the promotional framing suggests it might. Unsloth's separately published minimum-memory table pairs each bit-width with a floor — 223GB for 1-bit, 245GB for 2-bit, up to 810GB for Q8_0 — and the specific worked example, "the 2-bit dynamic quant UD-Q2_K_XL uses 254GB of disk space [and] works well in a 1×24GB GPU and 256GB of RAM with MoE offloading," is careful in a way that's easy to miss: it says 256GB of RAM, not "a 256GB Mac." That distinction matters, because the two aren't interchangeable — Apple's MacBook Pro line tops out at 128GB of unified memory even on the largest M-series chip, so any Mac actually offering 256GB has to be a Mac Studio, not a laptop. Unsloth doesn't make the MacBook claim here; it names a RAM figure and a GPU, and leaves the reader to supply the hardware. That's the more carefully worded of the two Mac-adjacent claims circulating around this release — a separate MLX-specific promotional claim, about running a smaller GLM-5.3-Flash variant on a MacBook Pro specifically, is covered in its own piece on this site and does not hold up the same way.

Between the two sections above, the pattern repeats: Z.ai's own claim about the model — same base, no architecture change — checks out exactly against the weights. Unsloth's claims about serving it check out on the totals and the hardware table, and come apart specifically where two adjacent SKUs get folded into one round number. Both are the kind of error that only shows up once the files exist to check against, which is the whole reason this update exists.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "GLM-5.3: the same base model, a month of post-training, and a cyber capability nobody planned", ai.thesatyajit.com, August 2026.

bibtex
@misc{ghana2026glm53,
  author = {Satyajit Ghana},
  title  = {GLM-5.3: the same base model, a month of post-training, and a cyber capability nobody planned},
  url    = {https://ai.thesatyajit.com/articles/glm-5-3},
  year   = {2026}
}
share