~/satyajit

Qwen3.8-27B variants: the bill is tokens times bytes

mdjsonmcp

2026-09-26 · 20 min · qwen · quantization · open-weights · inference-optimization · reasoning · efficiency · benchmarks

Between 20 and 24 September three derivatives of Qwen/Qwen3.8-27B landed on Hugging Face. All three start from the same dense 64-block hybrid: 48 Gated DeltaNet blocks, 16 full-attention blocks, a multi-token-prediction (MTP) head and a vision tower. I read that architecture out of the files in the Qwen3.8 open-weights article. All three are pitched as making that model better to live with, and they pull different levers.

For one user decoding one answer, a dense model is bound by memory bandwidth. Every generated token streams every decoder weight through the chip once. To first order, the time to an answer is

t≈N⋅BWt \approx \frac{N \cdot B}{W}

where NN is tokens per answer (thinking plus reply), BB is bytes read per generated token, and WW is memory bandwidth in bytes per second.

I read the Hub API, the configs and cards, and the safetensors headers by range request wherever a repository is not gated. Numbers are measured (computed from the files), reported (the publisher's) or reasoned (my arithmetic on someone else's numbers).

bottlecapai/ThinkingCap-Qwen3.8-27B@d2b4e7a · snapshot 2026-09-26
parameters
27.78B
repo size
55.59 GB
architecture
Qwen3_5ForConditionalGeneration
task
image-text-to-text
library
transformers
license
other
safetensors
18 shards
largest file
3.99 GB
files
47
downloads
532
likes
102
gated
auto
parameters by dtype
BF1627.78B
qwen3_8token-efficientefficient-thinking

repo last modified 2026-09-25

orcarouter/OrcaSAQ-2-27B@15d20d7 · snapshot 2026-09-26
parameters
6.77B
repo size
12.31 GB
quantizedQwen/Qwen3.8-27B
architecture
Qwen3_5ForCausalLM
task
text-generation
library
vllm
license
apache-2.0
safetensors
4 shards
largest file
3.99 GB
files
19
downloads
1.2K
likes
152
languages
en, zh
parameters by dtype
BF162.9MF1630.8MI81.27BI165.47B
qwenqwen3.8qwen3_5orcasaq2quantizationmixed-precision3-bitvllm

repo last modified 2026-09-24

Bytes per token, from the headers

BB is not the file size. Three things in the file are not read per token:

  1. The embedding table is a lookup. A token reads one row of it, 10 KB, not the 2.54 GB table.
  2. The vision tower does nothing for text.
  3. The MTP head is read only when you run speculative decoding.

What streams per decoded token is the linear layers of the 64 blocks, plus the LM head, plus a few tens of MB of norms and the narrow Gated DeltaNet gates. The decoder's linear layers are 24,326,963,200 weights in every build. The builds differ only in how many bits each weight gets. It is the same count-only-what-is-read rule that sets bits per weight in runtime dynamic compression.

builddecoder linears, bits per weightLM headBB per tokenon disklabel
Qwen3.8-27B, BF1616BF1651.25 GB55.56 GBmeasured
ThinkingCap FP88.00BF1626.93 GB31.24 GBreasoned (gated; from file sizes)
ThinkingCap MLX 4-bit DWQ6.048-bit19.76 GB22.54 GBmeasured
ThinkingCap NVFP44.50BF1616.28 GB20.59 GBmeasured
OrcaSAQ-23.216-bit10.79 GB12.27 GBmeasured

"On disk" is not like for like either. The BF16, FP8, NVFP4 and MLX builds carry the 0.92 GB vision tower; OrcaSAQ-2 does not.

Two rows get their own sections below: the "4-bit" MLX build is six bits per weight where it matters, and OrcaSAQ-2 moves 4.75× fewer bytes per token than BF16.

Before trusting the formula, check it against a measurement. BottleCap publishes single-request decode on one RTX PRO 6000 Blackwell, whose spec bandwidth is 1,792 GB/s. BF16 does 26.2 tok/s against a roofline of 35.0. FP8 does 44.7 against 66.5. NVFP4 does 68.0 against 110.1. So real decode reaches 62–75% of the floor, and the NVFP4 speedup it measured (2.6×) is smaller than the bytes ratio predicts (3.15×). The formula ranks builds correctly. It does not predict ratios to the decimal.

seconds per answer = tokens × bytes per token ÷ bandwidthbatch-1 decode roofline · bytes from headers · tokens from BottleCap
workload
hardware
ThinkingCap (NVFP4) on H200: 13,900 tokens × 16.28 GB ÷ 4,800 GB/s = 47.1 s
tokens lever, this workload
×0.797
bits lever, NVFP4
×0.318
bits lever, OrcaSAQ-2
×0.211
A floor, not a forecast: one request, weights only, no KV-cache reads, no prefill, no speculative decoding. The two levers multiply. At batch 1 the bits lever is the bigger one on all twelve of BottleCap's benchmarks; MMMLU, at ×0.346, comes closest to NVFP4's ×0.318. On a busy server the weights are shared across the batch and the ranking changes; see the measured table below.

Averaged over BottleCap's twelve benchmarks on an H200 (4,800 GB/s), the floor for the base in BF16 is 186.3 s per answer. ThinkingCap alone gets it to 148.4 s, ThinkingCap in NVFP4 to 47.1 s, and OrcaSAQ-2 to 39.2 s with no change to NN, an assumption, since OrcaRouter publishes no token counts.

At batch 1, bits are the bigger lever. ThinkingCap cuts mean total tokens from 17,450 to 13,900, a factor of 0.797. NVFP4 cuts bytes by a factor of 0.318 and OrcaSAQ-2 by 0.211. The token lever never catches up: its best case is MMMLU, 1,660 → 575 total tokens, a factor of 0.346.

The formula stops working on a busy server. There, weights are read once per step for the whole batch, and the cost moves to KV-cache reads and compute. Fewer weight bits help there only indirectly, by freeing memory for more cache. Fewer tokens cut the cache and the number of steps directly. BottleCap measured both levers on one axis: 32 MMLU-Pro questions, one H200, 16 concurrent requests, vLLM 0.26. Reported:

configmedian tokenss / task
Qwen3.8-27B, BF164848.9
Qwen3.8-27B, BF16, MTP5024.2
ThinkingCap, BF162323.7
ThinkingCap, BF16, MTP2162.0
ThinkingCap, NVFP42162.6
ThinkingCap, NVFP4, MTP2151.7

At 16 concurrent, the tokens bought 2.4× (8.9 → 3.7 s) and NVFP4 bought 1.4× on top (3.7 → 2.6 s). That is the reverse of the batch-1 ranking. BottleCap's throughput post goes further: on one RTX PRO 6000 with 8 GB of KV cache and 64 concurrent copies of one AIME prompt, traces 2.3× shorter finished 4.7× more tasks in 20 minutes, 112 against 24, because more short sequences fit in the cache at once. One prompt is not a workload, but the mechanism is sound.

The formula also ignores prefill. With 32,768-token prompts, the weight-only NVFP4 build (NVFP4A16, through the Marlin kernel) is no faster than BF16, 18.9 tok/s against 18.6, and its median first token comes later, 13,670 ms against 9,678. The FP8 build, which quantizes activations too, does 27.0 tok/s and 7,668 ms. The W4A4 build, with FP4 activations, does 32.8 tok/s and 6,393 ms, and is 29% slower than NVFP4 for a single request. Weight bits buy decode; activation bits buy prefill and batching.

Fewer tokens: ThinkingCap

What shipped

The BF16 checkpoint's 18 shards each have the same byte size as Qwen's, which is what the same tensors in the same dtype would give. It is gated (auto-approved) and licensed PolyForm Small Business 1.0.0 plus a personal-use grant, a real step down from the base's Apache-2.0. Five quantized builds sit alongside it.

How it was made

For this release, BottleCap does not say. The Qwen3.8 blog post states the objective ("Less number of thinking tokens to get the same answer") and one fact about training: the model "is trained at xhigh". The earlier Qwen3.6 post says a little more:

we trained on a curated set of problems covering multiple domains and difficulty levels. The training objective rewarded efficient reasoning rather than simply rewarding correctness.

"Rewarded" suggests reinforcement learning with a length-aware reward, reasoned from one word. No algorithm, dataset or length penalty is published for either model. The only other trace is the source path in the W4A4 build's quant_manifest.json, /shared/caveman_models/soup-r2-pubsoup. A directory name is not a method.

What the evaluation shows

The evaluation is the best part of the release. Both models ran through one harness on one H200 (vLLM 0.29.0), at reasoning_effort=xhigh, with Qwen's sampling (temperature 1.0, top_p 0.95, top_k 20) and MTP speculation, which BottleCap checked is accuracy-neutral: 98.23% on AIME without it, 98.13% with. Seeds run from 1 to 32, and every test set is complete except MMMLU, a fixed 10,000-question sample. Mean thinking tokens, reported:

benchmarkaccuracy, base → oursΔthinking tokens, base → ourscut
AIME 202698.13 → 94.27−3.8515,663 → 10,93430.2%
HMMT Feb 202695.83 → 94.70−1.1423,211 → 18,09922.0%
HMMT Nov 202597.08 → 96.04−1.0414,443 → 10,03730.5%
GPQA-Diamond89.93 → 88.04−1.8912,772 → 7,26743.1%
LiveCodeBench v691.14 → 91.21+0.0728,395 → 22,64520.3%
MMLU-Pro85.54 → 84.67−0.863,725 → 1,59157.3%
MMMLU85.38 → 84.09−1.291,656 → 57165.5%
IFBench79.75 → 79.71−0.047,961 → 4,26646.4%
RealWorldQA83.25 → 82.34−0.92992 → 49250.4%
AA-LCR81.75 → 84.00+2.252,550 → 1,56538.6%
τ²-bench76.16 → 75.15−1.014,584 → 3,16830.9%
Terminal-Bench 2.175.84 → 75.28−0.5672,871 → 65,09210.7%
mean86.65 → 85.79−0.8615,735 → 12,14437.2%

The τ²-bench and Terminal-Bench rows count reasoning summed over whole multi-turn episodes.

"37% on average" is a mean of percentages. The card says so in its evaluation notes: the bottom row is "the mean of the per-benchmark reductions, not the ratio of the two token figures beside it". Take the ratio of the two token figures instead, 12,144 against 15,735, and the cut is 22.8% (reasoned). For total tokens, thinking plus answer, it is 17,450 → 13,900, or 20.3%.

ThinkingCap vs Qwen3.8-27B · mean thinking tokens at xhigh
−37.2%mean of the twelve per-benchmark reductions (the headline)
AIME 2026−30.2%8.3%
HMMT Feb 2026−22.0%8.3%
HMMT Nov 2025−30.5%8.3%
GPQA-Diamond−43.1%8.3%
LiveCodeBench v6−20.3%8.3%
MMLU-Pro−57.3%8.3%
MMMLU−65.5%8.3%
IFBench−46.4%8.3%
RealWorldQA−50.4%8.3%
AA-LCR−38.6%8.3%
τ²-bench−30.9%8.3%
Terminal-Bench 2.1−10.7%8.3%
grey: base · teal: ThinkingCap · common scalecutweight
Counted per benchmark, MMMLU's 65.5% cut and Terminal-Bench's 10.7% get the same vote. Counted per token, Terminal-Bench alone is 38.6% of the weight and pulls the figure down to 22.8%. BottleCap's card states which one its headline is.

Both averages are honest; they answer different questions. The 37.2% is what a typical benchmark sees. The 22.8% is closer to what a workload drawn evenly from these twelve would pay, and the gap has a cause. The longest-running rows shrink least. Terminal-Bench is 38.6% of the pooled thinking tokens and gives up only 10.7% of them.

"Up to 65%" is MMMLU, 65.5%. It comes from a single seed, at a cost of 1.29 points of accuracy. It is not "65% fewer tokens at equal accuracy". Every row in the table is measured at an equal setting, xhigh on both sides, and accuracy moves where it moves. Ten of twelve benchmarks stay within two points. AIME 2026 loses 3.85, and the two intervals do not overlap (98.13 ± 0.74 against 94.27 ± 1.47). AA-LCR gains 2.25.

The equal-accuracy comparison that does hold up is against the dial the base model already has. Qwen3.8's chat template takes reasoning_effort as xhigh, medium or low, and turning it down is free. BottleCap ran both models at all three settings:

Scatter plot titled Reasoning efforts comparison. The y axis is accuracy in percent, macro-averaged, from about 76 to 88; the x axis is the share of mean thinking tokens saved against Qwen3.8-27B at xhigh, from 0% to 60%. A dashed line marks the base model at xhigh as the reference, at about 86.6%. The filled circle for ThinkingCap at xhigh sits just below the line at about 37% saved. The base model's medium and low settings, hollow square and triangle, sit near 52% and 55% saved but around 77 to 78% accuracy. ThinkingCap's medium and low settings sit further right, near 60% and 62%, and about a point lower.
The effort dial against the fine-tune. Turning the base model down to medium saves 52.1% of mean thinking tokens and costs 9.16 points. ThinkingCap at xhigh saves 37.2% and costs 0.86. The xhigh points average twelve benchmarks and the others eleven, because Terminal-Bench ran only at xhigh; the x axis is the per-benchmark mean, not pooled tokens. (BottleCap AI blog post, Figure 4.)

That is the real result. The dial buys its savings steeply: 9.16 points for 52.1%. ThinkingCap takes 37.2% for 0.86. At medium and low the fine-tune sits 7–8 points further right than the base at the same setting, for about one more point of accuracy, so the two cut different things and stack.

The other view is a budget: cap the tokens a response may generate, count anything unfinished as wrong, and sweep the cap.

Line chart titled Accuracy vs. token budget. The x axis is the generation-token budget per response on a log scale from about 10 to one million; the y axis is accuracy from 0 to 100 percent, counting a response as correct only if it finished within budget. Two S-shaped curves rise together. The ThinkingCap curve sits to the left of the Qwen3.8-27B curve from about 100 tokens up to about 50,000 tokens, meaning it reaches a given accuracy with a smaller budget. Both flatten near 87% above about 100,000 tokens.
Accuracy against a hard cap on generated tokens, across the benchmarks. Under a tight cap the shorter model wins; with no cap both finish at 87%. (BottleCap AI blog post, Figure 3.)

Under a 16K-token cap per response, ThinkingCap is more accurate than the base. If your stack already truncates, that is the version of the claim that pays.

Two side effects, both reported: truncation (a trace that never closes) drops from 0.51% to 0.34%, and MTP draft acceptance is unchanged, 53% against the base's 54%.

Fewer bits: OrcaSAQ-2

What SAQ is

The card describes "a proprietary sensitivity-aware mixed-precision quantization system", hence SAQ, sensitivity-aware quantization, and says the methodology is "not currently disclosed". The files disclose more than the card. config.json says, in part:

"quantization_config": {
  "quant_method": "exl3", "version": "1.5.1", "bits": 3, "head_bits": 6,
  "calibration": { "rows": 250, "cols": 2048 }, "codebook": "mul1",
  "mtp_bits": 4, "head_quant": "exl3-6bit", "embed_quant": "int8",
  "bits_per_weight": 3.21
}

exl3 is ExLlamaV3's format. It is a streamlined QTIP (arXiv 2406.11235): incoherence processing (blockwise Hadamard transforms, with the per-channel suh and svh vectors stored beside every tensor), then trellis-coded quantization of 16×16 tiles (the .trellis int16 tensors). OrcaRouter's own serving repository, Continuum-AI-Corp/OrcaSAQ2-kernel, says it outright: "a QTIP-style trellis code with a searched mixed-precision allocation", with "the trellis kernels" coming "from exllamav3". The codec is public. What is OrcaRouter's is the search that decides which tensor gets how many bits, and that search is unpublished.

The allocation can be read from the headers anyway. A 16×16 tile holds 256 weights, so at KK bits it packs 16K16K int16 values. The trellis tensor's last dimension is therefore 16K16K: 48 means 3 bits, 56 means 3.5, 64 means 4. The half rates come from the mul1 codebook, according to a comment in the kernel repository. I read all four shard headers and checked every one of the 400 decoder projections against the per-tensor bits_per_weight in quantization_config.json. All 400 agree. (The model card above shows 6.77B parameters because the Hub counts stored elements, not weights: 5,465,702,400 of them are int16 trellis words.)

OrcaSAQ-2 · bits per decoder projection, blocks 0–63read from trellis shapes in the safetensors headers
0816243240485663GDN in_proj_qkvblock 0 · GDN in_proj_qkv · 3 bitsblock 1 · GDN in_proj_qkv · 3 bitsblock 2 · GDN in_proj_qkv · 3 bitsblock 4 · GDN in_proj_qkv · 3 bitsblock 5 · GDN in_proj_qkv · 3 bitsblock 6 · GDN in_proj_qkv · 3 bitsblock 8 · GDN in_proj_qkv · 3 bitsblock 9 · GDN in_proj_qkv · 3 bitsblock 10 · GDN in_proj_qkv · 3 bitsblock 12 · GDN in_proj_qkv · 3 bitsblock 13 · GDN in_proj_qkv · 3 bitsblock 14 · GDN in_proj_qkv · 3 bitsblock 16 · GDN in_proj_qkv · 3 bitsblock 17 · GDN in_proj_qkv · 3 bitsblock 18 · GDN in_proj_qkv · 3 bitsblock 20 · GDN in_proj_qkv · 3 bitsblock 21 · GDN in_proj_qkv · 3 bitsblock 22 · GDN in_proj_qkv · 3 bitsblock 24 · GDN in_proj_qkv · 3 bitsblock 25 · GDN in_proj_qkv · 3 bitsblock 26 · GDN in_proj_qkv · 3 bitsblock 28 · GDN in_proj_qkv · 3 bitsblock 29 · GDN in_proj_qkv · 3 bitsblock 30 · GDN in_proj_qkv · 3 bitsblock 32 · GDN in_proj_qkv · 3 bitsblock 33 · GDN in_proj_qkv · 3 bitsblock 34 · GDN in_proj_qkv · 3 bitsblock 36 · GDN in_proj_qkv · 3 bitsblock 37 · GDN in_proj_qkv · 3 bitsblock 38 · GDN in_proj_qkv · 3 bitsblock 40 · GDN in_proj_qkv · 3 bitsblock 41 · GDN in_proj_qkv · 3 bitsblock 42 · GDN in_proj_qkv · 3 bitsblock 44 · GDN in_proj_qkv · 3 bitsblock 45 · GDN in_proj_qkv · 3 bitsblock 46 · GDN in_proj_qkv · 3 bitsblock 48 · GDN in_proj_qkv · 3 bitsblock 49 · GDN in_proj_qkv · 3 bitsblock 50 · GDN in_proj_qkv · 3 bitsblock 52 · GDN in_proj_qkv · 3 bitsblock 53 · GDN in_proj_qkv · 3 bitsblock 54 · GDN in_proj_qkv · 3 bitsblock 56 · GDN in_proj_qkv · 3 bitsblock 57 · GDN in_proj_qkv · 3 bitsblock 58 · GDN in_proj_qkv · 3 bitsblock 60 · GDN in_proj_qkv · 3 bitsblock 61 · GDN in_proj_qkv · 3 bitsblock 62 · GDN in_proj_qkv · 3 bitsGDN in_proj_zblock 0 · GDN in_proj_z · 3 bitsblock 1 · GDN in_proj_z · 3 bitsblock 2 · GDN in_proj_z · 3 bitsblock 4 · GDN in_proj_z · 3 bitsblock 5 · GDN in_proj_z · 3 bitsblock 6 · GDN in_proj_z · 3 bitsblock 8 · GDN in_proj_z · 3 bitsblock 9 · GDN in_proj_z · 3 bitsblock 10 · GDN in_proj_z · 3 bitsblock 12 · GDN in_proj_z · 3 bitsblock 13 · GDN in_proj_z · 3 bitsblock 14 · GDN in_proj_z · 3 bitsblock 16 · GDN in_proj_z · 3 bitsblock 17 · GDN in_proj_z · 3 bitsblock 18 · GDN in_proj_z · 3 bitsblock 20 · GDN in_proj_z · 3 bitsblock 21 · GDN in_proj_z · 3.5 bitsblock 22 · GDN in_proj_z · 3.5 bitsblock 24 · GDN in_proj_z · 3.5 bitsblock 25 · GDN in_proj_z · 3.5 bitsblock 26 · GDN in_proj_z · 3.5 bitsblock 28 · GDN in_proj_z · 3.5 bitsblock 29 · GDN in_proj_z · 3.5 bitsblock 30 · GDN in_proj_z · 3.5 bitsblock 32 · GDN in_proj_z · 3.5 bitsblock 33 · GDN in_proj_z · 3.5 bitsblock 34 · GDN in_proj_z · 3.5 bitsblock 36 · GDN in_proj_z · 3.5 bitsblock 37 · GDN in_proj_z · 3.5 bitsblock 38 · GDN in_proj_z · 3.5 bitsblock 40 · GDN in_proj_z · 3.5 bitsblock 41 · GDN in_proj_z · 3.5 bitsblock 42 · GDN in_proj_z · 3 bitsblock 44 · GDN in_proj_z · 3 bitsblock 45 · GDN in_proj_z · 3 bitsblock 46 · GDN in_proj_z · 3 bitsblock 48 · GDN in_proj_z · 3 bitsblock 49 · GDN in_proj_z · 3 bitsblock 50 · GDN in_proj_z · 3 bitsblock 52 · GDN in_proj_z · 3 bitsblock 53 · GDN in_proj_z · 3 bitsblock 54 · GDN in_proj_z · 3 bitsblock 56 · GDN in_proj_z · 3 bitsblock 57 · GDN in_proj_z · 3 bitsblock 58 · GDN in_proj_z · 3 bitsblock 60 · GDN in_proj_z · 3 bitsblock 61 · GDN in_proj_z · 3 bitsblock 62 · GDN in_proj_z · 3 bitsGDN out_projblock 0 · GDN out_proj · 3 bitsblock 1 · GDN out_proj · 3 bitsblock 2 · GDN out_proj · 3 bitsblock 4 · GDN out_proj · 3 bitsblock 5 · GDN out_proj · 3 bitsblock 6 · GDN out_proj · 3 bitsblock 8 · GDN out_proj · 3 bitsblock 9 · GDN out_proj · 3 bitsblock 10 · GDN out_proj · 3 bitsblock 12 · GDN out_proj · 3 bitsblock 13 · GDN out_proj · 3 bitsblock 14 · GDN out_proj · 3 bitsblock 16 · GDN out_proj · 3 bitsblock 17 · GDN out_proj · 3 bitsblock 18 · GDN out_proj · 3 bitsblock 20 · GDN out_proj · 3 bitsblock 21 · GDN out_proj · 3.5 bitsblock 22 · GDN out_proj · 3.5 bitsblock 24 · GDN out_proj · 3.5 bitsblock 25 · GDN out_proj · 3.5 bitsblock 26 · GDN out_proj · 3.5 bitsblock 28 · GDN out_proj · 3.5 bitsblock 29 · GDN out_proj · 3.5 bitsblock 30 · GDN out_proj · 3.5 bitsblock 32 · GDN out_proj · 3.5 bitsblock 33 · GDN out_proj · 3.5 bitsblock 34 · GDN out_proj · 3.5 bitsblock 36 · GDN out_proj · 3.5 bitsblock 37 · GDN out_proj · 3.5 bitsblock 38 · GDN out_proj · 3.5 bitsblock 40 · GDN out_proj · 3.5 bitsblock 41 · GDN out_proj · 3.5 bitsblock 42 · GDN out_proj · 3.5 bitsblock 44 · GDN out_proj · 3.5 bitsblock 45 · GDN out_proj · 3.5 bitsblock 46 · GDN out_proj · 3.5 bitsblock 48 · GDN out_proj · 3.5 bitsblock 49 · GDN out_proj · 3.5 bitsblock 50 · GDN out_proj · 3.5 bitsblock 52 · GDN out_proj · 3.5 bitsblock 53 · GDN out_proj · 3.5 bitsblock 54 · GDN out_proj · 3.5 bitsblock 56 · GDN out_proj · 3.5 bitsblock 57 · GDN out_proj · 3.5 bitsblock 58 · GDN out_proj · 3.5 bitsblock 60 · GDN out_proj · 3.5 bitsblock 61 · GDN out_proj · 3.5 bitsblock 62 · GDN out_proj · 3.5 bitsattn q_projblock 3 · attn q_proj · 3.5 bitsblock 7 · attn q_proj · 3.5 bitsblock 11 · attn q_proj · 3.5 bitsblock 15 · attn q_proj · 3.5 bitsblock 19 · attn q_proj · 3.5 bitsblock 23 · attn q_proj · 3.5 bitsblock 27 · attn q_proj · 3.5 bitsblock 31 · attn q_proj · 3.5 bitsblock 35 · attn q_proj · 3 bitsblock 39 · attn q_proj · 3 bitsblock 43 · attn q_proj · 3 bitsblock 47 · attn q_proj · 3 bitsblock 51 · attn q_proj · 3 bitsblock 55 · attn q_proj · 3 bitsblock 59 · attn q_proj · 3 bitsblock 63 · attn q_proj · 3 bitsattn k_projblock 3 · attn k_proj · 3.5 bitsblock 7 · attn k_proj · 3.5 bitsblock 11 · attn k_proj · 3.5 bitsblock 15 · attn k_proj · 3.5 bitsblock 19 · attn k_proj · 3.5 bitsblock 23 · attn k_proj · 3.5 bitsblock 27 · attn k_proj · 3.5 bitsblock 31 · attn k_proj · 3.5 bitsblock 35 · attn k_proj · 3.5 bitsblock 39 · attn k_proj · 3.5 bitsblock 43 · attn k_proj · 3.5 bitsblock 47 · attn k_proj · 3.5 bitsblock 51 · attn k_proj · 3.5 bitsblock 55 · attn k_proj · 3.5 bitsblock 59 · attn k_proj · 3.5 bitsblock 63 · attn k_proj · 3.5 bitsattn v_projblock 3 · attn v_proj · 4 bitsblock 7 · attn v_proj · 4 bitsblock 11 · attn v_proj · 4 bitsblock 15 · attn v_proj · 4 bitsblock 19 · attn v_proj · 4 bitsblock 23 · attn v_proj · 4 bitsblock 27 · attn v_proj · 4 bitsblock 31 · attn v_proj · 4 bitsblock 35 · attn v_proj · 4 bitsblock 39 · attn v_proj · 4 bitsblock 43 · attn v_proj · 4 bitsblock 47 · attn v_proj · 4 bitsblock 51 · attn v_proj · 4 bitsblock 55 · attn v_proj · 4 bitsblock 59 · attn v_proj · 4 bitsblock 63 · attn v_proj · 4 bitsattn o_projblock 3 · attn o_proj · 3 bitsblock 7 · attn o_proj · 3 bitsblock 11 · attn o_proj · 3 bitsblock 15 · attn o_proj · 3 bitsblock 19 · attn o_proj · 3 bitsblock 23 · attn o_proj · 3 bitsblock 27 · attn o_proj · 3 bitsblock 31 · attn o_proj · 3 bitsblock 35 · attn o_proj · 3 bitsblock 39 · attn o_proj · 3 bitsblock 43 · attn o_proj · 3 bitsblock 47 · attn o_proj · 3 bitsblock 51 · attn o_proj · 3 bitsblock 55 · attn o_proj · 3 bitsblock 59 · attn o_proj · 3 bitsblock 63 · attn o_proj · 3 bitsMLP gate_projblock 0 · MLP gate_proj · 3.5 bitsblock 1 · MLP gate_proj · 3.5 bitsblock 2 · MLP gate_proj · 3.5 bitsblock 3 · MLP gate_proj · 3.5 bitsblock 4 · MLP gate_proj · 3.5 bitsblock 5 · MLP gate_proj · 3.5 bitsblock 6 · MLP gate_proj · 3 bitsblock 7 · MLP gate_proj · 3 bitsblock 8 · MLP gate_proj · 3 bitsblock 9 · MLP gate_proj · 3 bitsblock 10 · MLP gate_proj · 3 bitsblock 11 · MLP gate_proj · 3 bitsblock 12 · MLP gate_proj · 2 bitsblock 13 · MLP gate_proj · 2 bitsblock 14 · MLP gate_proj · 2 bitsblock 15 · MLP gate_proj · 2 bitsblock 16 · MLP gate_proj · 2 bitsblock 17 · MLP gate_proj · 2 bitsblock 18 · MLP gate_proj · 3 bitsblock 19 · MLP gate_proj · 3 bitsblock 20 · MLP gate_proj · 3 bitsblock 21 · MLP gate_proj · 3 bitsblock 22 · MLP gate_proj · 3 bitsblock 23 · MLP gate_proj · 3 bitsblock 24 · MLP gate_proj · 3.5 bitsblock 25 · MLP gate_proj · 3.5 bitsblock 26 · MLP gate_proj · 3.5 bitsblock 27 · MLP gate_proj · 3.5 bitsblock 28 · MLP gate_proj · 3.5 bitsblock 29 · MLP gate_proj · 3.5 bitsblock 30 · MLP gate_proj · 3 bitsblock 31 · MLP gate_proj · 3 bitsblock 32 · MLP gate_proj · 3 bitsblock 33 · MLP gate_proj · 3 bitsblock 34 · MLP gate_proj · 3 bitsblock 35 · MLP gate_proj · 3 bitsblock 36 · MLP gate_proj · 3 bitsblock 37 · MLP gate_proj · 3 bitsblock 38 · MLP gate_proj · 3 bitsblock 39 · MLP gate_proj · 3 bitsblock 40 · MLP gate_proj · 3 bitsblock 41 · MLP gate_proj · 3 bitsblock 42 · MLP gate_proj · 3 bitsblock 43 · MLP gate_proj · 3 bitsblock 44 · MLP gate_proj · 3 bitsblock 45 · MLP gate_proj · 3 bitsblock 46 · MLP gate_proj · 3 bitsblock 47 · MLP gate_proj · 3.5 bitsblock 48 · MLP gate_proj · 3.5 bitsblock 49 · MLP gate_proj · 3.5 bitsblock 50 · MLP gate_proj · 3.5 bitsblock 51 · MLP gate_proj · 3.5 bitsblock 52 · MLP gate_proj · 3.5 bitsblock 53 · MLP gate_proj · 4 bitsblock 54 · MLP gate_proj · 4 bitsblock 55 · MLP gate_proj · 4 bitsblock 56 · MLP gate_proj · 4 bitsblock 57 · MLP gate_proj · 4 bitsblock 58 · MLP gate_proj · 4 bitsblock 59 · MLP gate_proj · 4 bitsblock 60 · MLP gate_proj · 4 bitsblock 61 · MLP gate_proj · 4 bitsblock 62 · MLP gate_proj · 4 bitsblock 63 · MLP gate_proj · 4 bitsMLP up_projblock 0 · MLP up_proj · 3 bitsblock 1 · MLP up_proj · 3 bitsblock 2 · MLP up_proj · 3 bitsblock 3 · MLP up_proj · 3 bitsblock 4 · MLP up_proj · 3 bitsblock 5 · MLP up_proj · 3 bitsblock 6 · MLP up_proj · 3 bitsblock 7 · MLP up_proj · 3 bitsblock 8 · MLP up_proj · 3 bitsblock 9 · MLP up_proj · 3 bitsblock 10 · MLP up_proj · 3 bitsblock 11 · MLP up_proj · 3 bitsblock 12 · MLP up_proj · 3 bitsblock 13 · MLP up_proj · 3 bitsblock 14 · MLP up_proj · 3 bitsblock 15 · MLP up_proj · 3 bitsblock 16 · MLP up_proj · 3 bitsblock 17 · MLP up_proj · 3 bitsblock 18 · MLP up_proj · 3 bitsblock 19 · MLP up_proj · 3 bitsblock 20 · MLP up_proj · 3 bitsblock 21 · MLP up_proj · 3 bitsblock 22 · MLP up_proj · 3 bitsblock 23 · MLP up_proj · 3 bitsblock 24 · MLP up_proj · 3.5 bitsblock 25 · MLP up_proj · 3.5 bitsblock 26 · MLP up_proj · 3.5 bitsblock 27 · MLP up_proj · 3.5 bitsblock 28 · MLP up_proj · 3.5 bitsblock 29 · MLP up_proj · 3.5 bitsblock 30 · MLP up_proj · 3 bitsblock 31 · MLP up_proj · 3 bitsblock 32 · MLP up_proj · 3 bitsblock 33 · MLP up_proj · 3 bitsblock 34 · MLP up_proj · 3 bitsblock 35 · MLP up_proj · 3 bitsblock 36 · MLP up_proj · 3 bitsblock 37 · MLP up_proj · 3 bitsblock 38 · MLP up_proj · 3 bitsblock 39 · MLP up_proj · 3 bitsblock 40 · MLP up_proj · 3 bitsblock 41 · MLP up_proj · 3.5 bitsblock 42 · MLP up_proj · 3.5 bitsblock 43 · MLP up_proj · 3.5 bitsblock 44 · MLP up_proj · 3.5 bitsblock 45 · MLP up_proj · 3.5 bitsblock 46 · MLP up_proj · 3.5 bitsblock 47 · MLP up_proj · 3 bitsblock 48 · MLP up_proj · 3 bitsblock 49 · MLP up_proj · 3 bitsblock 50 · MLP up_proj · 3 bitsblock 51 · MLP up_proj · 3 bitsblock 52 · MLP up_proj · 3 bitsblock 53 · MLP up_proj · 3.5 bitsblock 54 · MLP up_proj · 3.5 bitsblock 55 · MLP up_proj · 3.5 bitsblock 56 · MLP up_proj · 3.5 bitsblock 57 · MLP up_proj · 3.5 bitsblock 58 · MLP up_proj · 3.5 bitsblock 59 · MLP up_proj · 4 bitsblock 60 · MLP up_proj · 4 bitsblock 61 · MLP up_proj · 4 bitsblock 62 · MLP up_proj · 4 bitsblock 63 · MLP up_proj · 4 bitsMLP down_projblock 0 · MLP down_proj · 3 bitsblock 1 · MLP down_proj · 3 bitsblock 2 · MLP down_proj · 3 bitsblock 3 · MLP down_proj · 3 bitsblock 4 · MLP down_proj · 3 bitsblock 5 · MLP down_proj · 3 bitsblock 6 · MLP down_proj · 3 bitsblock 7 · MLP down_proj · 3 bitsblock 8 · MLP down_proj · 3 bitsblock 9 · MLP down_proj · 3 bitsblock 10 · MLP down_proj · 3 bitsblock 11 · MLP down_proj · 3 bitsblock 12 · MLP down_proj · 3 bitsblock 13 · MLP down_proj · 3 bitsblock 14 · MLP down_proj · 3 bitsblock 15 · MLP down_proj · 3 bitsblock 16 · MLP down_proj · 3 bitsblock 17 · MLP down_proj · 3 bitsblock 18 · MLP down_proj · 3 bitsblock 19 · MLP down_proj · 3 bitsblock 20 · MLP down_proj · 3 bitsblock 21 · MLP down_proj · 3 bitsblock 22 · MLP down_proj · 3 bitsblock 23 · MLP down_proj · 3 bitsblock 24 · MLP down_proj · 3 bitsblock 25 · MLP down_proj · 3 bitsblock 26 · MLP down_proj · 3 bitsblock 27 · MLP down_proj · 3 bitsblock 28 · MLP down_proj · 3 bitsblock 29 · MLP down_proj · 3 bitsblock 30 · MLP down_proj · 3 bitsblock 31 · MLP down_proj · 3 bitsblock 32 · MLP down_proj · 3 bitsblock 33 · MLP down_proj · 3 bitsblock 34 · MLP down_proj · 3 bitsblock 35 · MLP down_proj · 3.5 bitsblock 36 · MLP down_proj · 3.5 bitsblock 37 · MLP down_proj · 3.5 bitsblock 38 · MLP down_proj · 3.5 bitsblock 39 · MLP down_proj · 3.5 bitsblock 40 · MLP down_proj · 3.5 bitsblock 41 · MLP down_proj · 3 bitsblock 42 · MLP down_proj · 3 bitsblock 43 · MLP down_proj · 3 bitsblock 44 · MLP down_proj · 3 bitsblock 45 · MLP down_proj · 3 bitsblock 46 · MLP down_proj · 3 bitsblock 47 · MLP down_proj · 3.5 bitsblock 48 · MLP down_proj · 3.5 bitsblock 49 · MLP down_proj · 3.5 bitsblock 50 · MLP down_proj · 3.5 bitsblock 51 · MLP down_proj · 3.5 bitsblock 52 · MLP down_proj · 3.5 bitsblock 53 · MLP down_proj · 4 bitsblock 54 · MLP down_proj · 4 bitsblock 55 · MLP down_proj · 4 bitsblock 56 · MLP down_proj · 4 bitsblock 57 · MLP down_proj · 4 bitsblock 58 · MLP down_proj · 4 bitsblock 59 · MLP down_proj · 4 bitsblock 60 · MLP down_proj · 4 bitsblock 61 · MLP down_proj · 4 bitsblock 62 · MLP down_proj · 4 bitsblock 63 · MLP down_proj · 4 bitsblock mean33.5block 0: 3.12 bits per weightblock 1: 3.12 bits per weightblock 2: 3.12 bits per weightblock 3: 3.23 bits per weightblock 4: 3.12 bits per weightblock 5: 3.12 bits per weightblock 6: 3.00 bits per weightblock 7: 3.11 bits per weightblock 8: 3.00 bits per weightblock 9: 3.00 bits per weightblock 10: 3.00 bits per weightblock 11: 3.11 bits per weightblock 12: 2.77 bits per weightblock 13: 2.77 bits per weightblock 14: 2.77 bits per weightblock 15: 2.87 bits per weightblock 16: 2.77 bits per weightblock 17: 2.77 bits per weightblock 18: 3.00 bits per weightblock 19: 3.11 bits per weightblock 20: 3.00 bits per weightblock 21: 3.08 bits per weightblock 22: 3.08 bits per weightblock 23: 3.11 bits per weightblock 24: 3.32 bits per weightblock 25: 3.32 bits per weightblock 26: 3.32 bits per weightblock 27: 3.35 bits per weightblock 28: 3.32 bits per weightblock 29: 3.32 bits per weightblock 30: 3.08 bits per weightblock 31: 3.11 bits per weightblock 32: 3.08 bits per weightblock 33: 3.08 bits per weightblock 34: 3.08 bits per weightblock 35: 3.14 bits per weightblock 36: 3.20 bits per weightblock 37: 3.20 bits per weightblock 38: 3.20 bits per weightblock 39: 3.14 bits per weightblock 40: 3.20 bits per weightblock 41: 3.20 bits per weightblock 42: 3.16 bits per weightblock 43: 3.14 bits per weightblock 44: 3.16 bits per weightblock 45: 3.16 bits per weightblock 46: 3.16 bits per weightblock 47: 3.26 bits per weightblock 48: 3.27 bits per weightblock 49: 3.27 bits per weightblock 50: 3.27 bits per weightblock 51: 3.26 bits per weightblock 52: 3.27 bits per weightblock 53: 3.62 bits per weightblock 54: 3.62 bits per weightblock 55: 3.62 bits per weightblock 56: 3.62 bits per weightblock 57: 3.62 bits per weightblock 58: 3.62 bits per weightblock 59: 3.74 bits per weightblock 60: 3.74 bits per weightblock 61: 3.74 bits per weightblock 62: 3.74 bits per weightblock 63: 3.74 bits per weighty axis from 2.5 to 4 bits · outside the blocks: embedding int8, LM head 6-bit, MTP head 4-bit
2 bits3 bits4 bits3.5 bits
Two kinds of structure. By role, fixed everywhere: every value projection 4 bits, every key projection 3.5, every attention output and every GDN qkv projection 3. By depth, bands of about six blocks: the only 2-bit tensors are the MLP gates of blocks 12–17, and the MLPs of blocks 53–63 get 3.5 and 4. Block 0 gets nothing its neighbours do not.

Measured, over the 400 decoder projections:

Penjing-27B, another Qwen3.8-27B quant claiming a sensitivity-driven allocation, turned out to use a fixed first-and-last rule, identical at three bit budgets. OrcaSAQ-2 is not that. It spends its spare bits at the back of the network, cuts one tensor type in one band of the middle, and treats attention by role, which is what a search would plausibly produce. Without the search's output I cannot check that it was one.

Outside the blocks, the embedding is INT8 with a per-row scale (1.27 GB), the LM head is 6-bit EXL3 (954 MB), and the MTP head is 4-bit (213 MB). There is no vision tower: OrcaSAQ-2 is text-only.

The headline numbers

The launch post reads "55.59 → 12.06 GB — 78.3% smaller / 4.61×, 3.21 bpw · 93.2% Top-1 agreement". Each number, checked:

claimwhat it isverdict
3.21 bpw3.2114, over the 24,326,963,200 decoder linear weights onlyholds exactly, for that scope
whole checkpoint12,270,184,036 bytes over 27,320,697,856 parameters = 3.59 bpwnot claimed
55.59 GBevery file in Qwen/Qwen3.8-27B (55,586,114,863 bytes), including the vision tower, the MTP head and the tokenizer filesmatches
12.06 GBOrcaSAQ-2's tensors minus its own MTP head (12,057,547,332 bytes)matches
4.61×, 78.3%55.59 / 12.06right, but not like for like

All measured. The 12.06 is the only split of the files I could find that produces that number; OrcaRouter does not say how it was computed. Count the same things on both sides, text weights plus MTP head, and 54.64 GB → 12.27 GB is 4.45×, 77.5% smaller. The model card's own figures, 54 GB → 12.3 GB and "4.4× smaller", are the like-for-like version. The launch post took the larger numerator from one count and the smaller denominator from another.

93.2% top-1 agreement is token-level agreement with BF16's top choice over 16,376 predicted WikiText-2 tokens, from the same pass as the perplexity. One predicted token in about fifteen differs from what BF16 would have picked. The mean KL divergence is 0.031. Perplexity goes from 5.6468 to 5.6482, or +0.02%, and that is the weakest of the three: errors in both directions cancel inside a perplexity. That is how it can move by 0.02% while 6.8% of the argmaxes flip. These are fidelity numbers, not benchmark scores, and they are the right kind to publish.

70.0 on SWE-bench Verified and 58.4 on Terminal-Bench 2.1 are placed on the card next to Claude, Gemini and GPT scores, not next to Qwen3.8-27B in BF16. The harness is not stated. Three parties have published BF16 numbers for this exact base on Terminal-Bench 2.1:

sourceharnessTerminal-Bench 2.1SWE-bench Verified
Qwen's model cardTerminus73.0not reported
PrismML, in the Bonsai 2 whitepaperTerminus-2 under Harbor; mini-swe-agent69.780.6
BottleCap's ThinkingCap cardTerminus-2 under Harbor, 4 seeds75.84not reported
OrcaSAQ-2not stated58.470.0

58.4 sits 11.3 to 17.4 points under every published BF16 number, and 70.0 sits 10.6 under PrismML's. I cannot attribute that gap to quantization, because I do not know the harness. But same-harness BF16 against the quant is the comparison a quantization release exists to make, and it is the one missing. The card's own limitations list says as much: "Long-horizon comparisons should use a controlled same-harness evaluation."

OrcaRouter's two documents also disagree on serving. The card reports 90.1 tok/s single-stream with MTP under a 15.7 GiB cap; the kernel repository reports 111.2 tok/s and a KV pool of 35,617 tokens against the card's 14,563, and its last commit reads "the pool figure came from a 16K run, not the 32K it ships". Neither names the physical card, so the roofline check is not possible.

Fewer bits on fewer tokens: ThinkingCap's own builds

BottleCap quantized its own fine-tune five ways. Four of the five repositories are not gated, so those headers are readable:

buildhowon diskdecoder linearsevaluated against BF16
FP8block-wise E4M3, 128×12831.24 GB8.00 bpw (reasoned)5 benchmarks, paired
NVFP4NVFP4A16, llm-compressor20.59 GB4.50 bpw (measured)5 benchmarks, paired
NVFP4 W4A4, AWQ168 MLP layers FP4, 233 layers FP823.42 GBmixed5 benchmarks, paired; Blackwell only
MLX 4-bit DWQaffine, group 64, 4/8-bit mix22.54 GB6.04 bpw (measured)5 benchmarks
GGUFIQ4_XS / Q4_K_M / Q6_K / Q8_015.48 / 17.44 / 23.86 / 29.05 GB4.53 / 5.11 / 6.99 / 8.51 bpw, whole file5 benchmarks

The MLX "4-bit" build is 6.04 bits per weight on the decoder's linear layers. The card is candid: only the MLP projections of the first 56 blocks are 4-bit, and self-attention, the wide Gated DeltaNet projections, the last eight blocks' MLPs, the LM head and the embeddings are 8-bit. With group-64 affine scales and biases, 4-bit costs 4.5 bits and 8-bit costs 8.5, and only 61.6% of the decoder's weights are in the 4-bit set. The headers sum to 18,360,565,760 bytes of decoder linears, exactly what that layout predicts. It is a deliberate trade, but the label hides a consequence: this build reads 19.76 GB per token, 21% more than NVFP4's 16.28 GB. The launch post's "21 GB, down from 52 GB" is GiB on both sides (20.99 against 51.75), so it is consistent.

Quantizing also moved NN a little. On NVFP4, mean completion tokens fell 9.5% on IFBench and 12.2% on AA-LCR, with both intervals excluding zero, while the medians rose 7–9%: fewer very long answers, not uniformly shorter ones. NVFP4 accuracy stays within 1.7 points of the BF16 fine-tune on four benchmarks and loses 3.0 on AA-LCR, every interval including zero. MLX lands within about two points either way. All reported, on 100 to 1,500 questions per benchmark.

Different words: Hemmingway-1

Altworld/Hemmingway-1@b987f18 · snapshot 2026-09-26
parameters
26.90B
repo size
54.66 GB
architecture
Qwen3_5ForCausalLM
task
text-generation
library
transformers
license
cc-by-nc-4.0
safetensors
13 shards
largest file
4.99 GB
files
28
downloads
5.6K
likes
701
languages
en
parameters by dtype
BF1627.32B

These add to 27.32B, not the 26.90B total — expected when a packed format stores more than one value per element.

qwen3.8chatcreative-writingaltworld

repo last modified 2026-09-22

Hemmingway-1 is the control group. It is BF16 and text-only (Qwen3_5ForCausalLM, vision tower dropped). Its headers list exactly the base's 866 text and MTP tensors, with the same shapes and dtype, renamed from model.language_model.* to model.*. The chat template is byte-identical to the base's, so it thinks at xhigh by default. BB is unchanged at 51.25 GB per token (measured), and nobody has published NN.

The card's "Code" link is a GitHub repository with a README, the model card and six chart images; there is no code and no word on data or method. The license is CC BY-NC 4.0, stricter than the base's. The claims are about style, on three internal benchmarks that the card flags up front: CommunicationBench is 80 real requests judged pairwise and blind, in both orders, by an outside model. There it scores 1026 against Fable 5.1's 1024, a two-point lead on 80 requests. On EQ-Bench 4, the one public benchmark, it places third at 1330, and the base model is not on the chart. One chart does bear on cost:

Bar chart titled The Message, Not a Memo, labeled internal, subtitled Replies wrapped in commentary or options, lower is better. Bars in ascending order: GPT-6 Astra 4, Qwen3.8 27B base 19, Grok 4.6 22, Hemmingway 1 39 (highlighted), Fable 5.1 67, Kimi K3 92, Fable 5 93, GLM-5.3 93.
Share of replies that bury the requested text in commentary or options. The card's prose names the models above nine in ten; its own chart shows the base model at 19 and the fine-tune at 39. (Altworld, Hemmingway-1 model card, Figure 4.)

The fine-tune wraps its answer twice as often as the model it came from, 39 against 19. More wrapper probably means more answer tokens (reasoned; no token counts are published). Hemmingway-1 changes what the tokens say, not how many there are or what each one costs. Others pulled the bits lever for it: a Hub search for "Hemmingway-1" returns 52 repositories besides Altworld's own, most of them quantized builds.

What nobody measured

The practical reading is short. For one user on one card, bytes per token is the number to shop on, and OrcaSAQ-2's 10.79 GB is the smallest here. For a full server, tokens per answer is the number, and ThinkingCap is the only release that moved it, with seeds and intervals to show for it. BottleCap is also the only publisher that measured both levers, and in its own table the build that pulls both is the fastest.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Qwen3.8-27B variants: the bill is tokens times bytes", ai.thesatyajit.com, September 2026.

bibtex
@misc{ghana2026qwen3827bvariants,
  author = {Satyajit Ghana},
  title  = {Qwen3.8-27B variants: the bill is tokens times bytes},
  url    = {https://ai.thesatyajit.com/articles/qwen3-8-27b-variants},
  year   = {2026}
}
share