~/satyajit

Penjing-27B: reading a quantisation claim out of the GGUF header

mdjsonmcp

2026-09-22 · 15 min · quantization · gguf · llama-cpp · inference · local-llm · measurement

A GGUF release is one of the few artefacts in this field that cannot lie to you. The filename is marketing; the tensor-info block at the top of the file is the build. Every tensor's name, shape and ggml type id is sitting there in the first few megabytes, and a Range: bytes=0-8388607 request is enough to read the whole allocation without downloading 6 GB of weights.

So when PollardWeights/Penjing-27B-Pollard appeared today with a ladder of quants of Qwen/Qwen3.8-27B and the line "a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush)", the useful thing to do was not to read the card. It was to read the headers.

PollardWeights/Penjing-27B-Pollard@6439855 · snapshot 2026-09-22
repo size
35.02 GB
quantizedQwen/Qwen3.8-27B
task
image-text-to-text
library
trellis
license
apache-2.0
gguf files
5
largest file
12.94 GB
files
8
downloads
4
likes
4
languages
en
pollard-weightspollardggufllama.cppik_llama.cpptrellisimatrixqwen3_5

repo last modified 2026-09-22

This article is about the method, not the model. Qwen3.8-27B has been covered elsewhere; what is new here is Pollard Weights, the toolkit that produced these files, and whether "measured allocation" names something specific enough to check. It does, and it checks out partly.

What "Pollard" is

The repo never says, so this is a reading rather than a fact: pollarding is the horticultural practice of cutting a tree back to the same knuckles year after year so it keeps its shape at a size you choose, which fits a tool whose pitch is "tell it the target's budget and it fits that." Not a person, not Pollard's rho. The model here is called Penjing, the Chinese art the Japanese word bonsai descends from — and this site has already covered a ternary Bonsai built on the same base model, which turns out to be the most useful comparison available. More on that below.

WestWaters/pollard-weights@62fdbf1 · snapshot 2026-09-22
tracked files
498
license
Apache-2.0
branch
main
tests
22 files
source
2.2 MB
commit date
2026-09-21
source by language
Python1.9 MB(132)JavaScript136.9 kB(1)CSS53.6 kB(1)C++16.6 kB(1)Shell12.2 kB(6)HTML9.4 kB(1)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-09-22 at 62fdbf1 branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, testFileCount

The mechanism, per the toolkit's own README, is a two-signal pipeline:

  1. pollard-sensitivity points at a model and measures which tensor groups matter, by crushing each group one at a time and watching the KL against f16, plus that model's own noise curve.
  2. pollard-fit --sensitivity then allocates bits to minimise KL for a stated size budget — --ram/--vram of the target machine, minus a working reserve, rather than a bit-width chart.

The README's claimed wins for that procedure are +6–27% lower KL on a dense 1.5B (5 of 5 sizes) and +21–43% on a 40-expert MoE (4 of 5), against "uniform imatrix-IQ at matched size", measured on 48K held-out wikitext tokens. Here is its own chart:

Two line charts side by side, both with file size in gigabytes on the x-axis and mean KL divergence against f16 on a log y-axis, labelled 'lower is better'. Left panel, Dense, Qwen2.5-1.5B: a blue 'pollard' line and a grey 'uniform imatrix-IQ' line run nearly on top of each other from about 0.56 GB down to 1.27 GB, with green annotations reading +21%, +6%, +14%, +27%, +15% at five matched sizes. Right panel, MoE, granite-3B-a800m with 40 experts: the blue line sits clearly below the grey one from 1.4 GB onward, annotated +29%, +31%, +43%, +21%, with a red −15% marking the one loss at the smallest size.
Pollard Weights' own evidence for measured allocation — on a 1.5B dense model and a 40-expert MoE, not on Penjing-27B. Lower KL is better; the annotations are relative KL reductions at matched file size. (Pollard Weights, README, assets/benchmarks/kl_win.png.)

Two things about that figure matter for what follows. It is real evidence, published with the script that regenerates it from raw data — python experiments/plot_kl_win.py — which puts it ahead of most quantisation marketing. And it is not evidence about this release: different models, different sizes, a different architecture class. The README is careful about this in a comment aimed at readers like me:

DENSE models -> imatrix-guided K-quants (IQ3_S / IQ4_XS / Q6_K). The importance matrix is the dense win; the per-layer measured-KL knapsack does NOT pay off on dense (no expert redundancy to reallocate from), so we don't claim it there. MoE models -> pollard-automap measured expert-allocation. Rule of thumb: imatrix = dense, automap = MoE.

Qwen3.8-27B is dense. config.json has no num_experts field at all: 64 blocks, hidden 5120, intermediate 17408, 24 query heads over 4 KV heads, and a hybrid attention stack — three gated-delta linear-attention blocks for every one full-attention block (full_attention_interval: 4). By the toolkit's own rule of thumb, this release is on the imatrix path, and the per-layer knapsack is the leg the toolkit says does not pay off here.

So what does the per-layer structure in these files actually look like?

The allocation, read off the files

Four of the six advertised rungs were in the repo when I checked at 15:46 UTC on 2026-09-22. I pulled the header of each — Range requests against huggingface.co/.../resolve/main/<file>.gguf, parsing magic, version, tensor count, the metadata KV block and then the tensor-info array — and resolved the ggml type ids, including ik_llama.cpp's extended range (153 = IQ2_KT, 158 = IQ1_KT), against ik_llama.cpp's own ggml.h. Every file reports the same 866 tensors and the same 27,320,697,856 elements.

tensor types per block, read from the GGUF header7.36 GB · 2.150 bpw measured
ffn_downMLP write-back — never crushedblock 0 · ffn_down · IQ2_KT (2.125 bpw)block 1 · ffn_down · IQ2_KT (2.125 bpw)block 2 · ffn_down · IQ2_KT (2.125 bpw)block 3 · ffn_down · IQ2_KT (2.125 bpw)block 4 · ffn_down · IQ2_KT (2.125 bpw)block 5 · ffn_down · IQ2_KT (2.125 bpw)block 6 · ffn_down · IQ2_KT (2.125 bpw)block 7 · ffn_down · IQ2_KT (2.125 bpw)block 8 · ffn_down · IQ2_KT (2.125 bpw)block 9 · ffn_down · IQ2_KT (2.125 bpw)block 10 · ffn_down · IQ2_KT (2.125 bpw)block 11 · ffn_down · IQ2_KT (2.125 bpw)block 12 · ffn_down · IQ2_KT (2.125 bpw)block 13 · ffn_down · IQ2_KT (2.125 bpw)block 14 · ffn_down · IQ2_KT (2.125 bpw)block 15 · ffn_down · IQ2_KT (2.125 bpw)block 16 · ffn_down · IQ2_KT (2.125 bpw)block 17 · ffn_down · IQ2_KT (2.125 bpw)block 18 · ffn_down · IQ2_KT (2.125 bpw)block 19 · ffn_down · IQ2_KT (2.125 bpw)block 20 · ffn_down · IQ2_KT (2.125 bpw)block 21 · ffn_down · IQ2_KT (2.125 bpw)block 22 · ffn_down · IQ2_KT (2.125 bpw)block 23 · ffn_down · IQ2_KT (2.125 bpw)block 24 · ffn_down · IQ2_KT (2.125 bpw)block 25 · ffn_down · IQ2_KT (2.125 bpw)block 26 · ffn_down · IQ2_KT (2.125 bpw)block 27 · ffn_down · IQ2_KT (2.125 bpw)block 28 · ffn_down · IQ2_KT (2.125 bpw)block 29 · ffn_down · IQ2_KT (2.125 bpw)block 30 · ffn_down · IQ2_KT (2.125 bpw)block 31 · ffn_down · IQ2_KT (2.125 bpw)block 32 · ffn_down · IQ2_KT (2.125 bpw)block 33 · ffn_down · IQ2_KT (2.125 bpw)block 34 · ffn_down · IQ2_KT (2.125 bpw)block 35 · ffn_down · IQ2_KT (2.125 bpw)block 36 · ffn_down · IQ2_KT (2.125 bpw)block 37 · ffn_down · IQ2_KT (2.125 bpw)block 38 · ffn_down · IQ2_KT (2.125 bpw)block 39 · ffn_down · IQ2_KT (2.125 bpw)block 40 · ffn_down · IQ2_KT (2.125 bpw)block 41 · ffn_down · IQ2_KT (2.125 bpw)block 42 · ffn_down · IQ2_KT (2.125 bpw)block 43 · ffn_down · IQ2_KT (2.125 bpw)block 44 · ffn_down · IQ2_KT (2.125 bpw)block 45 · ffn_down · IQ2_KT (2.125 bpw)block 46 · ffn_down · IQ2_KT (2.125 bpw)block 47 · ffn_down · IQ2_KT (2.125 bpw)block 48 · ffn_down · IQ2_KT (2.125 bpw)block 49 · ffn_down · IQ2_KT (2.125 bpw)block 50 · ffn_down · IQ2_KT (2.125 bpw)block 51 · ffn_down · IQ2_KT (2.125 bpw)block 52 · ffn_down · IQ2_KT (2.125 bpw)block 53 · ffn_down · IQ2_KT (2.125 bpw)block 54 · ffn_down · IQ2_KT (2.125 bpw)block 55 · ffn_down · IQ2_KT (2.125 bpw)block 56 · ffn_down · IQ2_KT (2.125 bpw)block 57 · ffn_down · IQ2_KT (2.125 bpw)block 58 · ffn_down · IQ2_KT (2.125 bpw)block 59 · ffn_down · IQ2_KT (2.125 bpw)block 60 · ffn_down · IQ2_KT (2.125 bpw)block 61 · ffn_down · IQ2_KT (2.125 bpw)block 62 · ffn_down · IQ2_KT (2.125 bpw)block 63 · ffn_down · IQ2_KT (2.125 bpw)block 64 · ffn_down · Q6_K (6.5625 bpw)ffn_gate · ffn_upMLP body — crushed on 61 of 64 blocksblock 0 · ffn_gate · ffn_up · IQ2_KT (2.125 bpw)block 1 · ffn_gate · ffn_up · IQ2_KT (2.125 bpw)block 2 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 3 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 4 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 5 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 6 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 7 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 8 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 9 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 10 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 11 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 12 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 13 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 14 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 15 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 16 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 17 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 18 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 19 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 20 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 21 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 22 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 23 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 24 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 25 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 26 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 27 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 28 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 29 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 30 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 31 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 32 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 33 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 34 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 35 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 36 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 37 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 38 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 39 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 40 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 41 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 42 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 43 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 44 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 45 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 46 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 47 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 48 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 49 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 50 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 51 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 52 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 53 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 54 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 55 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 56 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 57 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 58 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 59 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 60 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 61 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 62 · ffn_gate · ffn_up · IQ1_S (1.5625 bpw)block 63 · ffn_gate · ffn_up · IQ2_KT (2.125 bpw)block 64 · ffn_gate · ffn_up · Q6_K (6.5625 bpw)attn_q · attn_outputfull-attention blocks only — never crushedblock 3 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 7 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 11 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 15 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 19 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 23 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 27 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 31 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 35 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 39 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 43 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 47 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 51 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 55 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 59 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 63 · attn_q · attn_output · IQ2_KT (2.125 bpw)block 64 · attn_q · attn_output · Q6_K (6.5625 bpw)attn_k · attn_vcrushed on 15 of 16 — saves 11 MBblock 3 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 7 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 11 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 15 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 19 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 23 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 27 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 31 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 35 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 39 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 43 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 47 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 51 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 55 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 59 · attn_k · attn_v · IQ1_S (1.5625 bpw)block 63 · attn_k · attn_v · IQ2_KT (2.125 bpw)block 64 · attn_k · attn_v · Q6_K (6.5625 bpw)attn_qkv · attn_gate · ssm_*linear-attention blocks only — never crushedblock 0 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 1 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 2 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 4 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 5 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 6 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 8 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 9 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 10 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 12 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 13 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 14 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 16 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 17 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 18 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 20 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 21 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 22 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 24 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 25 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 26 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 28 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 29 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 30 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 32 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 33 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 34 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 36 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 37 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 38 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 40 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 41 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 42 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 44 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 45 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 46 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 48 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 49 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 50 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 52 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 53 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 54 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 56 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 57 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 58 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 60 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 61 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)block 62 · attn_qkv · attn_gate · ssm_* · IQ2_KT (2.125 bpw)0116324863mtptransformer block index
IQ1_S1.5625 bpwIQ2_KT2.125 bpwQ6_K6.5625 bpwtensor absent on this block
Nominal bits-per-weight are the ggml block sizes (IQ1_S 50 B/256, IQ2_XXS 66, IQ3_S 110, Q4_K 144, Q6_K 210; ik_llama.cpp's IQ1_KT 56 and IQ2_KT 68). The measured file bpw in the header above is lower or higher than any single row because token_embd and output sit outside this grid.

The three low rungs have exactly one piece of per-layer structure, and it is the same piece in all three: blocks 0, 1, 63 and 64 keep the rung's primary type; blocks 2 through 62 have ffn_gate and ffn_up demoted to IQ1_S. Not a gradient, not a knapsack boundary that moves with the budget — one binary split, at the same four indices, at 1.91 bpw and at 2.12 bpw and at 2.15 bpw. Block 64 is not a transformer block at all: qwen35.nextn_predict_layers is 1, so it is the multi-token-prediction head, and it is kept at Q6_K (or Q2_K in the smallest build).

A KL-minimising allocation re-solved under three different size budgets would not normally produce the identical keep-set three times. A first-and-last-blocks heuristic would, and first-and-last is the oldest protective rule in llama.cpp's own mix.

The IQ3_S rung has no per-layer structure whatsoever. All 64 blocks get IQ3_S on ffn_down, ffn_gate and ffn_up; the only departures from uniform are by roleattn_v and attn_qkv bumped to Q4_K, token_embd to Q4_K, output to Q6_K — and those are the shape llama.cpp produces by itself for MOSTLY_IQ3_S on a model whose n_gqa is at least 4. Qwen3.8-27B's is 6. If there is a Pollard-specific allocation in the IQ3_S file, it is not visible in the tensor types.

The one role-level choice that is genuinely its own

There is a departure from stock, and it runs the wrong way. In the three low rungs attn_k and attn_v are demoted to IQ1_S on 15 of the 16 full-attention blocks, while attn_q and attn_output keep the primary type. llama.cpp does the opposite: for every IQ2-class file type it bumps attn_v to Q4_K when n_gqa >= 4, on the standard argument that K and V are tiny and sensitive.

Tiny is the operative word. On this model attn_k and attn_v are 5120 × 1024 each — 4 KV heads at head_dim 256 — 5.24M parameters per copy, so 78.6M across the 15 demoted blocks for attn_k and the same again for attn_v. Dropping each group from 2.125 bpw to 1.5625 bpw saves 5.5 MB, 11 MB out of a 7.34 GB file: 0.15%. The same demotion applied to ffn_gate and ffn_up over 61 blocks moves 10.87B parameters and saves 765 MB, 10.4% of the file. One of those is the size lever. The other is a rounding error bought with the most sensitivity-critical tensors in the stack.

That is the kind of thing a per-tensor KL sweep is supposed to prevent, and it is the single clearest place where I would want to see the sweep's output rather than the claim.

Nothing weighs what its name says

where the bytes went, by tensor rolesummed from GGUF tensor shapes × block sizes
IQ1_KT6.36 GiB · 1.906 bpw
ffn_gate + ffn_up: 2299 MB (35.3% of the file)ffn_down: 1277 MB (19.6% of the file)linear attention (48 blocks): 1224 MB (18.8% of the file)full attention (16 blocks): 398 MB (6.1% of the file)output.weight (LM head): 874 MB (13.4% of the file)token_embd.weight: 417 MB (6.4% of the file)MTP head + norms: 20 MB (0.3% of the file)
IQ2_XXS7.07 GiB · 2.120 bpw
ffn_gate + ffn_up: 2408 MB (33.3% of the file)ffn_down: 1544 MB (21.3% of the file)linear attention (48 blocks): 1441 MB (19.9% of the file)full attention (16 blocks): 509 MB (7.0% of the file)output.weight (LM head): 874 MB (12.1% of the file)token_embd.weight: 417 MB (5.8% of the file)MTP head + norms: 46 MB (0.6% of the file)
IQ2_KT7.17 GiB · 2.150 bpw
ffn_gate + ffn_up: 2412 MB (32.8% of the file)ffn_down: 1588 MB (21.6% of the file)linear attention (48 blocks): 1485 MB (20.2% of the file)full attention (16 blocks): 521 MB (7.1% of the file)output.weight (LM head): 874 MB (11.9% of the file)token_embd.weight: 417 MB (5.7% of the file)MTP head + norms: 46 MB (0.6% of the file)
IQ3_S12.62 GiB · 3.785 bpw
ffn_gate + ffn_up: 5048 MB (39.1% of the file)ffn_down: 2524 MB (19.5% of the file)linear attention (48 blocks): 2731 MB (21.1% of the file)full attention (16 blocks): 818 MB (6.3% of the file)output.weight (LM head): 1043 MB (8.1% of the file)token_embd.weight: 715 MB (5.5% of the file)MTP head + norms: 46 MB (0.4% of the file)
ffn_gate + ffn_upffn_downlinear attention (48 blocks)full attention (16 blocks)output.weight (LM head)token_embd.weightMTP head + norms
Bars are to a common scale, so IQ3_S reads as twice the file rather than the same bar restacked. In the 6.53 GB IQ1_KT build the LM head is 874 MB and the embedding table 417 MB — 19.8% of the file, in two tensors that the rung's name does not describe.

Measured bits-per-weight, computed from the same headers — element count times the ggml block size, summed, divided by the parameter count:

filenominal typemeasured bpwsizeruns in
Penjing-27B-IQ1_KT.gguf1.751.9066.53 GBik_llama.cpp
Penjing-27B-IQ2_XXS.gguf2.06252.1207.25 GBany llama.cpp
Penjing-27B-IQ2_KT.gguf2.1252.1507.36 GBik_llama.cpp
Penjing-27B-IQ3_S.gguf3.43753.78512.94 GBany llama.cpp

The gap is almost entirely two tensors. output.weight — the LM head, 1.27B parameters over a 248,320-token vocabulary — is Q5_K in all three low rungs, 874 MB, 13.4% of the 6.53 GB file. token_embd.weight is Q2_K, another 417 MB. Together they are 19.8% of the smallest build, at an average of 4.0 bits per weight, inside a file whose name says 1.75.

Protecting the LM head at low bit is correct and I would do the same. The point is narrower: a reader picking a rung by its name is not picking the precision of a fifth of the file, and the card's "which file should I choose?" section — which is written in RAM budgets, sensibly — does not say so either.

Two smaller things fall out of the same headers. The imatrix metadata is preserved in every file: quantize.imatrix.entries_count 496, chunks_count 510, and quantize.imatrix.dataset reading C:\pollard\phome\gemma4_calib.txt. The card describes the calibration corpus as "a mixed-domain corpus so the matrix sees every register the model serves"; the artefact says it is a calibration file carried over from a Gemma-4 build. Those are compatible — calibration text is text, and reusing a good corpus across models is normal practice — but the filename is the more informative of the two descriptions, and it is the one that ships inside the file. The builds were also made on Windows, which is why the IQ1_KT build's header is 18 MB and the IQ3_S build's is 11 MB.

The part that deserves straightforward credit

quality against size, with the holes left in3 of 6 rungs measured · 4 of 6 uploaded
44.555.56681216243256file size (GB, log scale)lower is betterf16 · 54.64 GBIQ1_KT · 6.527 GB · 1.906 bpw · perplexity 5.7864IQ1_KT1.906 bpwIQ2_XXS · 7.25 GB · 2.12 bpw · perplexity 5.4312IQ2_XXS2.12 bpwIQ2_KT · 7.361 GB · 2.15 bpw · perplexity 5.0817IQ2_KT2.15 bpwIQ3_Sfile, no numberIQ4_XSno fileQ6_Kno file
Quality figures are the card's (wikitext2, ctx 2048, single machine). Sizes and bits-per-weight are measured from the published GGUF headers on 2026-09-22. IQ4_XS and Q6_K are placed at the sizes the card states; no such file was in the repo at that time.

Most GGUF repos publish a size column and nothing else. This one publishes perplexity against an f16 reference, mean KL divergence, and top-1 agreement — the share of tokens where the quantised model's argmax matches what f16 would have picked, which is the number that actually predicts whether a build feels broken. llama-perplexity prints it as Same top p and almost nobody reports it.

The numbers are also not flattering, which is the tell that they were measured rather than chosen:

rungPPLvs f16mean KLDtop-1 agree
f16 reference4.1422
IQ1_KT5.7864+39.7%0.551375.88%
IQ2_XXS5.4312+31.1%0.514276.38%
IQ2_KT5.0817+22.7%0.421779.10%

Read the top-1 column rather than the perplexity one. At the recommended default rung, one token in five is not the token f16 would have produced. A mean KL of 0.42 nats is a large distributional change. The card says as much implicitly by calling IQ2_KT "recommended default" rather than "near-lossless", and its errata line — "Single machine; replication invited" — is the right register. But an 87%-smaller file that disagrees with the original on 21% of tokens is a different product from a compression, and the honest framing is that this ladder is for people who want a 27B-class model on a 10 GB machine at all, not for people who want the 27B model in less space.

What 2.12 bpw costs this model on actual tasks

Perplexity and top-1 agreement are cheap to measure and weakly related to whether a model can still do a job. For this particular base model there happens to be a better answer already on the site. PrismML's Bonsai 2 whitepaper evaluates Qwen3.8-27B at IQ2_XXS — the same nominal quant type as Penjing's second rung, on the same base model — across 20 benchmarks in thinking mode, and I walked through that table in September. The relevant column:

Qwen3.8-27B FP16Qwen3.8-27B IQ2_XXSloss
GPQA Diamond90.5165.45−25.1
LiveCodeBench v690.0570.05−20.0
AIME2694.5878.60−16.0
τ²-Bench (agentic)82.7369.43−13.3
MMLU-Redux91.4685.79−5.7
GSM8K97.1995.38−1.8
Overall (20)85.475.2−10.2

That is a different IQ2_XXS file — different builder, different imatrix, almost certainly a different mix — so it is a proxy and not a measurement of Penjing's build. But it is the same model at the same nominal type and near enough the same size, and it says something the perplexity table cannot: at this bit budget the damage is not spread evenly. GSM8K barely notices. GPQA Diamond loses a quarter of its score. A "+31% perplexity" summary and a 25-point GPQA drop are the same file.

This is also why Bonsai is the right foil. It lands at 1.72 bits per weight and 5.95 GB on the same base — smaller than Penjing's 1.91 bpw / 6.53 GB rung — and claims 83.9 on the same 20-benchmark aggregate against FP16's 85.4. The difference is not a better allocator. It is quantisation-aware training: Bonsai retrained the model to live in ternary weights, and Penjing is post-training quantisation of a checkpoint that was never asked. Post-training quantisation at under two bits and QAT at under two bits are different products that happen to produce files of the same size, and no allocator closes that gap.

Also worth stating plainly: IQ3_S, IQ4_XS and Q6_K have no published quality number, and as of 15:46 UTC on 2026-09-22 IQ4_XS and Q6_K had no published file — the card's "which file should I choose?" section recommends Q6_K for stock llama.cpp, and llama-server -hf PollardWeights/Penjing-27B-Pollard:Q6_K does not resolve. The repo was mid-upload: the four GGUFs landed between 05:44 and 13:54 UTC the same day. That is a timing observation, not a criticism; by the time you read this the ladder may be complete. It is in the article because the three rungs with numbers are the three smallest, and the shape of the quality curve above 2.15 bpw is exactly the region a buyer cares about and nobody has measured.

What I would run next

All of it is cheap and none of it needs a new idea.

  1. The sensitivity sweep's output, as a file. pollard-sensitivity produces a per-tensor KL ranking. Ship it next to the GGUFs. It converts "bits placed by per-layer sensitivity" from a claim into an artefact, and it would immediately settle whether the 0/1/63/64 keep-set is what the sweep said or what the default said.
  2. An ablation on attn_k/attn_v. Build one more IQ2_KT with K and V left at the primary type. It costs 11 MB. If top-1 agreement moves by more than a point, the demotion is a bad trade at any budget; if it does not, that is a genuinely interesting result about hybrid-attention models and worth a note of its own.
  3. PPL and top-1 for the upper rungs, which is one llama-perplexity run per file.
  4. The same three numbers for stock llama.cpp IQ2_XXS at matched size, produced with the same imatrix. That is the comparison the toolkit's own chart makes on other models and the one this repo does not make on this one.

What would change my mind

6 claims above, and what would falsify each

  1. The per-layer structure in the low rungs is a fixed first-and-last rule, not a per-budget KL knapsack.

    A fourth low rung at a different size budget whose protected set is not {0, 1, 63, 64}. Or the pollard-sensitivity output for this model showing blocks 0, 1 and 63 ranked top and 2–62 genuinely flat — in which case the rule and the measurement agree and my inference from the keep-set being identical across three budgets is simply wrong. My evidence is the tensor types in four files on one day; it cannot distinguish a measurement that landed on the obvious answer from a default that was never measured.

  2. The IQ3_S rung shows no Pollard-specific allocation.

    A tensor-type listing from a stock llama.cpp --imatrix IQ3_S build of the same model that differs from this file. I am inferring "stock" from the fact that every departure from uniform IQ3_S matches a rule in llama-quant.cpp for n_gqa >= 4; I did not run the baseline build, and token_embd at Q4_K in particular could equally be an explicit --token-embedding-type.

  3. Demoting attn_k and attn_v buys 11 MB and costs sensitivity.

    The 11 MB is arithmetic on the shapes and is not in doubt. The cost is: an IQ2_KT variant with K and V at the primary type that scores no better on top-1 agreement than the shipped one would show the demotion is free on this architecture — plausible, since only 16 of 64 blocks use full attention at all.

  4. Nineteen point eight percent of the 6.53 GB file is the embedding table and the LM head.

    A recount from the header. output.weight is 248,320 × 5120 at Q5_K (176 B per 256) = 874.1 MB; token_embd.weight is the same shape at Q2_K (84 B per 256) = 417.2 MB; the tensor-byte total is 6,509 MB. If the repo republishes the file with a different LM-head type, this number moves.

  5. PrismML's IQ2_XXS column is a usable proxy for what this bit budget costs on tasks.

    A 20-benchmark run of Penjing's own IQ2_XXS that lands materially above 75.2 overall, or above 65.45 on GPQA Diamond. It plausibly could: Pollard pins ffn_down at the primary type and protects the LM head at Q5_K, and PrismML's build may do neither. If Penjing's file scores meaningfully better at the same size, that is the strongest possible evidence for measured allocation and it belongs on the model card.

  6. Only three of the six advertised rungs carry a quality measurement.

    A card update. This was true at 15:46 UTC on 2026-09-22, four hours after the last upload, and the repo was visibly still being filled — the README gained a top-1 column at 05:36 the same morning. A later commit adding PPL for IQ3_S, IQ4_XS and Q6_K would overturn it, and would be the single most useful thing the repo could add.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Penjing-27B: reading a quantisation claim out of the GGUF header", ai.thesatyajit.com, September 2026.

bibtex
@misc{ghana2026penjing27b,
  author = {Satyajit Ghana},
  title  = {Penjing-27B: reading a quantisation claim out of the GGUF header},
  url    = {https://ai.thesatyajit.com/articles/penjing-27b},
  year   = {2026}
}
share