# Mach-2 Additive Medium: 1.70 bits, a trellis, and a 102 GB table the bits leave out

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mach-2-additive
> date: 2026-10-06
> tags: quantization, mixture-of-experts, qwen, llama-cpp, gguf, benchmarks

This site has now covered five different ways of shrinking [Qwen3.8-Flash-Next](/articles/qwen3-8-flash-next), and every one of them had to decide what to do with the model's 51B-parameter n-gram embedding table. So when [Syzygy announced](https://x.com/syzygyeng/status/2107537110937010387) a build "at 1.7 bits per weight on RAM", "9.4x smaller than its base model", that beats Qwen3.8-27B on DeepSWE, NL2Repo, Terminal-Bench and AutomationBench, the first thing I did was add up the files.

They come to 129 GB.

That isn't a contradiction, and the model card is upfront about why. But the number is a good way in. Mach-2 Additive Medium is a careful piece of engineering wrapped in two pieces of framing. "Additive" names a codec family it isn't really a member of, and "1.7 bits on RAM" describes the part of the model that sits in VRAM. Under the framing is a QTIP-style trellis quantizer with per-expert bit rates, which I think is the most interesting thing Syzygy shipped and the thing the launch never mentions.

<ModelCard
  repo="SyzygyResearch/Mach-2-Additive-Medium"
  claimed="1.70 bits per weight, 26.7 GB for 125.7B text parameters"
  note="The 26.7 GB is the text weights only. The repo also carries the n-gram embedding table at BF16 (22 shards, 102.4 GB) and a 1.5 GB 4-bit MTP draft for MLX. The GGUF repo is one 130.1 GB file holding both."
/>

## 129 GB, of which 26.7 are the 1.7 bits

The native repo is a set of packed safetensors files plus a `decode.py` that turns them back into a BF16 checkpoint. I read the header of every one of the 132 `.safetensors` files with HTTP range requests and summed what they hold. The text model is 125,743,653,795 parameters in 26,717,083,816 bytes of tensors, which is **1.6998 bits per weight**. `MANIFEST.json` says 125,743,653,760 and 26,717,634,752 bytes, the difference being file headers and a few int64 offsets. So the headline bit rate is real, to four digits.

Then there is `packed/table/`: 22 files, 128 BF16 tensors of shape `[2500012, 160]`, holding 51,200,245,760 parameters in 102,400,491,520 bytes. It is the n-gram embedding table from the base model, the trigram lookup the [Flash-Next technical report](/articles/qwen3-8-flash-next) spends a whole section justifying. Syzygy did not compress it at all. The card says so in its first sentence, "plus the model's n-gram embedding table at bf16 (102.4 GB)", and the GGUF's README tells you to `--mlock` it into about 105 GB of free host RAM.

So the 9.4x needs its denominator stated. BF16 text weights are 125.74B × 2 bytes = 251.5 GB, and 251.5 / 26.7 = 9.41. Correct. Put the table on both sides, since the BF16 base ships it too, and the ratio is 353.9 GB against 129.1 GB, or **2.74x**. Averaged over all 176.9B parameters the files hold, the bit rate is 5.84.

Which is the right number depends on what you're asking. The table is read 16 rows per token, 5,120 bytes, so for decode bandwidth it barely exists. For "will it fit on my machine", it is four fifths of the download. "1.7 bits per weight on RAM" is the odd one: following Syzygy's own instructions, the 1.7-bit part goes to the GPU and the BF16 part is what sits in RAM.

<BitLedger />

The ledger also shows where the bits actually go. The routed experts are 120.8B of the 125.7B text parameters, 96.1%, and they average **1.565 bits**. Everything else costs four bits or more. The attention, linear-attention and shared-expert matrices (2.91B parameters) are at 4.005 bits, the LM head at 5.25 (int5, group 64), the token embedding at 4.5 (int4 with an FP16 min and max per 64), the hyper-connection mixers at about 8, and 78.4M router, norm and recurrent-scalar parameters stay at BF16. That's the usual shape of an extreme quant: sub-2-bit numbers are only possible on a mixture of experts, because there the experts are almost the whole model.

The experts don't all get the same rate, either. Each expert's three matrices carry a `rate_k4` byte, bits per group of four weights, and the shards bucket experts by it. Across all 48 layers and 24,576 experts:

| bits per weight | 1.0 | 1.25 | 1.5 | 1.75 | 2.0 |
|---|---:|---:|---:|---:|---:|
| experts | 3,628 | 6,530 | 2,774 | 3,278 | 8,366 |
| share | 14.8% | 26.6% | 11.3% | 13.3% | 34.0% |

A third of the experts get two bits and one in seven gets one. The allocation also leans by depth: layer 0's experts average 1.69 bits, and the last nine layers sit between 1.48 and 1.53. How Syzygy chose the rates isn't documented. The exporter isn't public, and the fork's own docs say so: "The exporter that produces the code streams is not part of this repo."

## A trellis wearing an additive name

"Additive" has a specific meaning in quantization, and I expected to find it. [AQLM](https://arxiv.org/abs/2401.06118) (Egiazarian et al., "Extreme Compression of Large Language Models via Additive Quantization") stores each group of eight weights as the **sum** of codewords drawn from several learned codebooks. Two codebooks of 256 entries, say, give 65,536 representable vectors for 16 bits of index. This is multi-codebook additive quantization. It sits in a line of vector quantizers next to [QuIP#](https://arxiv.org/abs/2402.04396), which rotates weights with a randomized Hadamard transform until they look Gaussian and then rounds groups of eight to a fixed E8 lattice codebook, and [VPTQ](https://arxiv.org/abs/2409.17066), which learns vector codebooks with second-order information.

Mach-2 doesn't do this. No tensor in either repo is a stack of codebooks whose entries get summed. What the files hold is the next step in that same line, trellis-coded quantization, and its closest published relative is [QTIP](https://arxiv.org/abs/2406.11235) ("Quantization with Trellises and Incoherence Processing", Tseng, Sun, Hou and De Sa).

### From one code per weight to one path per tile

The reason anyone bothers with any of this is the distortion you pay per bit. Most GGUF quants are scalar: each weight gets its own small integer and a group shares a scale. (The IQ2 and IQ3 types are the exception; they borrow lattice codebooks from QuIP#.) At two bits, a scalar quantizer has four levels per weight and no way to exploit the fact that weights come in groups. QTIP's Table 1 puts numbers on it for an i.i.d. Gaussian source at 2 bits: the optimal scalar (Lloyd-Max) quantizer reaches 0.118 mean squared error, QuIP#'s 8-dimensional E8P codebook 0.089, a 256-dimensional trellis code 0.069, against an information-theoretic floor of 0.063.

Vector quantization gets its gain by quantizing eight weights jointly, but a codebook for 8 dimensions at 2 bits has 2^16 entries, and going to 16 dimensions would need 2^32. A trellis gets most of the high-dimensional gain without that table. Instead of choosing a codeword per group, it chooses a **path** through a state machine for a whole sequence of weights, and each state emits a value. Viterbi search finds the path with the least squared error; storage is just the bits that steer it.

<Figure
  src="https://ai.thesatyajit.com/articles/mach-2-additive/fig5.png"
  alt="Left: a four-state trellis graph G with states 00, 01, 11 and 10. Right: six unquantized values 0.54, 0.03, 0.72, 0.19, 0.26, 0.89 above a trellis unrolled over six steps, with a red path chosen through it; each state maps to a codebook value (0.5, 0.1, 0.3, 0.8), giving quantized values 0.5, 0.1, 0.8, 0.1, 0.3, 0.8."
  caption="Trellis-coded quantization in miniature: the encoder picks the path through the trellis whose states' values best match the sequence, and only the path is stored (QTIP paper, Figure 2)."
/>

QTIP's practical contribution is the "bitshift" trellis. If each state is just the last L bits of the stream, the next state is the old one shifted left by k bits with k fresh bits on the end. Any state can be read straight off the bitstream at position t × k, so decoding is parallel and the trellis graph never has to be stored. The remaining cost is the state-to-value table, 2^L entries for L = 16, and QTIP's second contribution was to compute values from the state with a hash rather than look them all up.

<Figure
  src="https://ai.thesatyajit.com/articles/mach-2-additive/fig3.png"
  alt="Diagram titled QTIP. On the left, an incoherent weight matrix W divided into tiles of size Tx by Ty, with arrows labelled BlockLDLQ. On the right, a bitshift trellis quantizer: four columns of states 00, 01, 11, 10 with edges between them and a red path, labelled ultra-high dimensional quantization."
  caption="QTIP quantizes each tile of a Hadamard-rotated (incoherent) weight matrix as one long trellis path, which makes the effective dimension the whole tile (QTIP paper, Figure 1)."
/>

<Figure
  src="https://ai.thesatyajit.com/articles/mach-2-additive/fig4.png"
  alt="Line chart titled Llama 2 Scaling: WikiText-2 perplexity at 4,096 context against total model size in GB on a log scale. QTIP 2-bit, 3-bit and 4-bit curves run below the dashed QuIP# 2-bit and AQLM 2-bit curves; the QTIP 2-bit curve nearly coincides with an FP16-at-4-bit reference line."
  caption="On Llama 2, QTIP's 2-bit models sit below QuIP# and AQLM at the same size, and close to what 4-bit would give. This is the paper's result on dense models, not a measurement of Mach-2 (QTIP paper, Figure 1)."
/>

### What the spine stores

The non-expert matrices are where the lineage is easiest to see, because they use QTIP's hybrid code almost verbatim. Each 16 × 16 tile is 64 int16 words, 1,024 bits for 256 weights, so 4 bits per weight. `decode.py` builds the value table like this:

```python title="decode.py:207-221 (SyzygyResearch/Mach-2-Additive-Medium)"
NE_K, NE_L, NE_V, NE_TLUT_BITS, NE_TD = 4, 16, 2, 9, 16

def _ne_full_lut(tlut):
    ...
    s = np.arange(1 << NE_L, dtype=np.int64)
    p = s * (s + 1)
    row = (p >> (16 - NE_TLUT_BITS - 1)) & ((1 << NE_TLUT_BITS) - 1)
    table = small[row].copy()
    table[:, 0] *= (1 - ((p >> 15) & 1) * 2).astype(np.float32)
```

Compare QTIP's Algorithm 3, "HYB": hash the state as x · x + x, use bits from position 15 − Q upward as an index into a 2^Q × 2 codebook, and flip a sign from bit 15. With Q = 9, that's the same code, down to which bits it takes. The 512-row table is `packed/ne/tlut.safetensors`, and when I loaded it every entry turned out to be an integer between −32 and 31 times one step of 0.0867, which is what the fork's `ggml.h` means when it calls this the "rotated int-lattice trellis". Each 16-bit window yields two weights, so the window advances 8 bits a step, and the last windows wrap around to the start of the tile. QTIP calls that tail-biting; it saves storing a separate start state.

The widget below runs this decoder on the first real tile of layer 0's shared-expert gate projection, using the shipped words and the shipped table. Its first two outputs, −1.560 and −0.693, are what `decode.py` produces for the same bytes before the tile's scale and rotations are undone.

<TrellisWalk />

Slide it and watch the window. Half of every state is the previous state's bits, which is the trellis constraint at work: consecutive pairs of weights are not free to be anything, they're linked through the shared bits, and the encoder's Viterbi search exploits that link to get below what independent rounding could do at the same rate.

### What the experts store

The routed experts use a different codec, which the GGUF metadata names: `{"codec": "d4", "L": 16, "V": 4, "rates_k4": [4, 5, 6, 7, 8]}`. Same bitshift idea, a 16-bit window over a tail-biting stream, but each state now emits four weights and the window advances K4 bits, where K4 is the expert's `rate_k4`. At K4 = 4 that's one bit per weight, at K4 = 8 two bits. `decode.py` reads it directly:

```python title="decode.py:114-121"
def _ex_states(words, K4):
    words = np.asarray(words).view(np.uint16).astype(np.int64)
    rows, nstep, step = words.shape[0], EX_TD * EX_TD // EX_V, int(K4)
    bits = ((words[:, :, None] >> np.arange(15, -1, -1)) & 1).reshape(rows, -1)[:, :step * nstep]
    bits = np.concatenate([bits, bits[:, :EX_L - step]], axis=1)
    w = 1 << np.arange(EX_L - 1, -1, -1, dtype=np.int64)
    idx = np.arange(nstep)[:, None] * step + np.arange(EX_L)[None, :]
    return bits[:, idx] @ w
```

The state indexes a 65,536 × 4 table instead of a hash, shipped in `packed/experts/codebook.safetensors` (one table for 1 bit, one for 2 bits, a shared one for the rates between). I downloaded the 3 MB file to look at the entries. They are also small integers times a step: for the 1.0 to 1.75-bit rates, every coordinate is 0, ±1, ±2 or ±4 steps, and only about 2,400 distinct four-vectors appear among the 65,536 states; the 2-bit table adds ±3 and ±6 and has 14,571 distinct vectors. "d4" is not the D4 lattice, by the way. D4 points have an even coordinate sum, and barely half of these do. I read the name as "four-dimensional".

Varying the rate per expert is the part I'd call new. QTIP's published models use one rate per model, and Tim Dettmers' [runtime dynamic compression](/articles/runtime-dynamic-compression) paper explicitly blamed QTIP's weak showing below 2 bits on fractional rates being impractical, since encoding a matrix "can take days of GPU time". Syzygy's answer is five rungs, each a different K4 on the same 16-bit trellis, assigned per expert. Whatever their exporter does to pick the rung, it ran over 24,576 experts.

### Rotations, signs from a hash, and tiles that share a scale

Trellis codes assume Gaussian-looking inputs, and real weight matrices have outliers. Mach-2 handles that the way QuIP# and QTIP do: each matrix is multiplied on both sides by a random-sign diagonal and a Hadamard matrix before quantizing, and decoding undoes it, `W = diag(sv) · H · Ŵ · H · diag(su)` in the fork's notation. Two details caught my eye.

The dimensions here aren't powers of two. An expert's matrices are 640 × 2560, and 2560 = 20 × 128 while 640 = 20 × 32. `decode.py` builds 12- and 20-point Hadamard matrices from quadratic residues (Paley's construction) and takes Kronecker products with ordinary power-of-two butterflies. The CUDA kernels hard-code 12 × 12 and 20 × 20 sign tables for the same job.

And in the native format, the expert sign vectors aren't stored. `sign_seed()` hashes a string like `canon|L12|e301|gate|SU` with SHA-256, and `expert_signs()` feeds the first seven bytes into what the constants identify as a Philox4x32-10 counter generator. 24,576 experts × 3 matrices × 2 sides of signs are regenerated at load. The GGUF stores them instead, as FP16, which costs 0.47 GB.

The last piece is `wave_gamma`: per expert and matrix, 200 FP16 scales, one per anti-diagonal of the 40 × 160 grid of tiles (199 diagonals; one of the 200 slots goes unused). Each tile's decoded values are multiplied by its diagonal's scale before the rotations. I don't know why anti-diagonals; it is cheap side information, 400 bytes per matrix against 410 KB of codes at 2 bits.

### So where is the "additive"?

Nowhere in the Mach-2 math. In the fork, "additive" is a label on an older payload: `src/models/mach1.cpp` describes format version 3 as "additive (V8 experts, rotated NE spine, int5-g64 head, nibble-LUT embed)", and version 2 added a shared low-rank correction to "demoted" experts, `y = B (c_e ⊙ (A x))`, which is genuinely additive in the residual sense. Mach-2 is format version 5 and uses neither. My best guess is that the name came along with the product line. A second possibility is Syzygy's own site, which advertises "Multiplication-Free Models at 1.7 Bits per Weight". Codebook coordinates of 0, ±1, ±2 and ±4 can be applied with shifts and adds. But that's my reading, not theirs, and the shipped kernels don't do it: the CPU path multiplies with int8 dot-product instructions, and the CUDA path with FP16 tensor cores. Mach-2 also has a 4-bit spine and an int5 head that are ordinary multiplies however you build them.

## Decoding a matrix you never materialise

A packed checkpoint is only useful if you can run it without first expanding it back to BF16, and the [llama.cpp-mach1 fork](https://github.com/SyzygyResearch/llama.cpp-mach1) is where that happens. The codec is not a ggml quant type. The fork adds custom ops (`GGML_OP_MACH1_D4_MM` for experts, `GGML_OP_MACH1_RT_MM` for the spine, `GGML_OP_MACH1_HEAD_MM` for the head) and stores the code streams under plain int16 and int8 tensor types. Stock llama.cpp can't load the file, and the docs warn that the fork's ggml libraries must never be mixed with stock builds, because "the op enum differs - an ABI break, not a graceful failure".

<RepoCard repo="SyzygyResearch/llama.cpp-mach1" />

The CPU reference makes the trick clear. For each (token, expert) pair, `ggml_compute_forward_mach1_d4_mm` (`ggml/src/ggml-cpu/ops.cpp`) multiplies the input by the expert's `su` signs and runs a Hadamard transform on it, so the input lives in the same rotated space as the codes. It then quantizes that rotated vector to int8 in groups of 16, walks each tile's trellis to get 64 four-vectors of 4-bit integers, multiplies the integer codes against the int8 activations with `dpbusd`-style instructions, and finally scales, rotates back and applies `sv`. The decoded matrix never exists in memory. The inner decode is three lines:

```cpp title="ggml/src/ggml-cpu/ops.cpp:13870-13875 (llama.cpp-mach1 @ 31261ba)"
for (int j = 0; j < 64; ++j) {
    const int b  = j*K4;
    const int by = b >> 3;
    const uint32_t w24 = ((uint32_t) buf[by] << 16) | ((uint32_t) buf[by + 1] << 8) | buf[by + 2];
    z[j] = z8[(w24 >> (8 - (b & 7))) & 0xFFFFu];
}
```

Read 24 bits at the step's bit offset, shift out the 16-bit state, look up four packed nibbles. On CUDA, the walk kernel decodes tiles into shared memory as FP16, split into a high and a low half to keep precision, and feeds them to tensor cores with `wmma`. The fork also pattern-matches whole regions of the decode graph into fused kernels (`gdn_full`, `shexp_gu`, `qkvzb`, `exp_mega`), and its docs tell you to check a kernel census with `GGML_MACH1_TIME=1`, because if those fusions don't engage, "decode will be much slower than it should be".

The honest line in the docs is about the CPU: "usable but slow relative to llama.cpp's hand-tuned quant kernels - roughly 3x behind Q4_K_M decode on AVX2 hosts and ~5x on AVX-512". A trellis walk costs integer work per weight that a Q4_K block doesn't, and on a CPU that is enough to lose despite reading far fewer bytes. The fork's verdict is "A GPU is strongly recommended", and for CUDA it's the only fast path; AMD builds with HIP "run the codec ops on the CPU". The fork's Vulkan shaders cover the Mach ops but aren't certified bit-identical to the CPU reference.

## Memory and speed, from the bytes

What follows is for the fork and the GGUF.

The GGUF holds more than the card's 26.7 GB. I parsed its header (2,405 tensors) and summed every tensor: 130.07 GB, of which the table is 102.40 and the rest **27.67 GB**. The extra gigabyte over the native format is mostly the stored sign vectors, the embedding in its lossless nibble-LUT form, and the router kept at F32. With `-ngl 99` those 27.67 GB go to the GPU, so the card's "32 GB or more" is right with a few gigabytes to spare.

The KV cache is small, because only 12 of the 48 layers are full attention, each with 2 KV heads of 256: 24,576 bytes per token at FP16, plus a little for the sparse-attention indexer's keys. That's 0.81 GB at a 32,768-token context and 6.44 GB at the full 262,144. The 36 Gated DeltaNet layers add a fixed recurrent state of about 113 MB at FP32. Long agent runs are where this model's architecture already helps, and the compression leaves that alone.

The table goes to host memory. The fork tags it as a layer input, the same placement llama.cpp gives token embeddings, so with `--mlock` you need about 105 GB of free RAM and `ulimit -l unlimited`. Without `--mlock` it is memory-mapped and pages the OS drops are read back from disk. At 16 rows of 320 bytes per token the bandwidth is nothing; the cost is latency on cold pages. So the practical box is a 32 GB GPU plus 128 GB of RAM and 130 GB of disk. Syzygy says they checked it on one A100 80 GB.

Speed is where I'd temper expectations, and the reason is in the third ledger view. Decoding one token touches 10 of 512 experts in each layer but every non-expert matrix, the head and the mixers. By my count that's 3.18 GB per token, and the experts are only 0.46 GB of it, 14.5%. The 4-bit spine alone is 1.46 GB. For comparison, the same touched parameters at a flat 4.5 bits would be 3.75 GB. So against a conventional 4-bit build, Mach-2 cuts the bytes per token by about 15%, while cutting the model's resident size by more than half. It is a capacity codec far more than a bandwidth codec, and the trellis decode work comes on top.

Dividing memory bandwidth by 3.18 GB gives a ceiling, not a prediction: about 86 tokens/s on an M4 Pro's 273 GB/s, 172 on an M4 Max's 546 GB/s, and 641 on an A100's 2,039 GB/s, with one token per step. The launch thread claims "over 110 tokens per second on a 48GB MacBook". That comes from Syzygy's own macOS engine, not the fork, and nothing names the chip, context or settings. On an M4 Pro it is above the single-token ceiling, which would need speculative decoding; the repo does ship a 4-bit copy of the base model's MTP layer "for speculative decoding with MLX". On an M4 Max it fits without. A 48 GB Mac also cannot hold the 102 GB table, so that run must page the table from SSD. I couldn't check the 110.

## Eight benchmarks, one real gap

The card's table compares Mach-2 with Qwen3.8-Flash-Next at BF16, "both run on the same harness, settings and tasks", and the footnotes are unusually specific: Terminal-Bench 2.1 with Terminus 2 and four hours per task; DeepSWE with Claude Code as the agent, up to 10.5 hours and 700 steps per task; NL2Repo in OpenHands without internet; τ³-Banking with GPT-5.4 mini as simulated user and judge. Running a 1.7-bit model through 10-hour agent episodes is expensive, and most quantization releases stop at perplexity. Credit for that.

<Figure
  src="https://ai.thesatyajit.com/articles/mach-2-additive/fig1.png"
  alt="Scatter chart of average score (HLE, Terminal-Bench 2.1, AutomationBench, DeepSWE) against GB of weights on GPU on a log scale. DeepSeek V4.1 Flash original at 307.5 GB scores 63.9; Mach DeepSeek V4.1 Flash at 2.0 bits and 146.8 GB scores 63.1, 98.7% of original. Qwen3.8 Flash-Next original BF16 at 251.4 GB scores 60.4; Mach-2 Additive Medium at 1.70 bits and 26.7 GB scores 58.2, 96.3% of original. Qwen3.8 27B original BF16 at 54.0 GB scores 51.5. Bonsai 2 27B at 1.76 bits and 5.9 GB scores 35.2, 68.4% of the 27B. A dashed line marks 80 GB, one H100."
  caption="Syzygy's headline chart. The x axis counts weights on the GPU, so the 102.4 GB n-gram table is not on it. The Qwen3.8-27B point is a four-benchmark average; per-benchmark 27B scores are not published (Syzygy, model card and launch post)."
/>

The averages check out. Mach-2's four scores average 58.215 and BF16's 60.43, a ratio of 96.3%, as plotted. The more useful thing I could do with the table was turn percentages back into task counts, because most of them land on whole numbers. Humanity's Last Exam is 731 against 860 correct out of 2,158. Terminal-Bench is 70 against 73 out of 89. DeepSWE is 55 against 56 out of 113. AIME is 231 against 230 answers out of 240.

<NoiseFloor />

Put a rough 95% interval around each difference and only one gap survives. HLE's 5.98-point drop is about four standard errors. Every agentic difference is a handful of tasks inside an interval of 5 to 14 points: a model that matched BF16 exactly would produce gaps this size. The interval is crude (I can't pair the runs, because per-task results aren't published, and pairing would tighten it), but it says the agentic rows are consistent with no loss, which is a weaker claim than the 98.2% DeepSWE retention suggests. AutomationBench at 102.0% and AIME at 100.4% are the same noise pointing the other way.

The one difference that is real sits on the knowledge exam. HLE is a single-turn test of facts and reasoning, exactly the kind of "short benchmark" the launch says traditional quants preserve. A plausible reading is that squeezing 96% of the parameters to about 1.5 bits costs stored knowledge first and agentic behaviour later, if at all. That fits what other heavily compressed models have shown, but this table is the evidence I have, and it is one benchmark.

The comparison with Qwen3.8-27B is harder to check. The post says Mach-2 outperforms it "on benchmarks such as DeepSWE v1.1, NL2Repo, Terminal Bench, and Automation Bench". What Syzygy published for the 27B is a single point on the chart: 51.5 averaged over HLE, Terminal-Bench, AutomationBench and DeepSWE. NL2Repo isn't in that average, and no per-benchmark 27B scores appear in the post, the thread or either card. A 6.7-point lead on a four-benchmark average is believable for a 125B-parameter MoE against a 27B dense model. The per-benchmark claim I couldn't verify.

## The long-horizon argument, and what backs it

The launch's thesis is the interesting sentence: "Traditional quantizations fall short. They preserve short benchmark performance but collapse over long horizons." It's a plausible mechanism. Small per-token errors compound over a long trajectory, a model that drifts loops or overthinks, and a 50-step agent run gives drift more chances than a single answer does.

The evidence offered is one chart, posted in the thread with the line "A strong heuristic for long horizon performance retention is response length, typically increased by overthinking."

<Figure
  src="https://ai.thesatyajit.com/articles/mach-2-additive/fig2.png"
  alt="Chart titled Answer length vs the original on AIME, subtitle Same 60 problems, 8 answers each, shaded bars are 95% intervals. Mach-2 Additive Medium at 1.70 bits produces 106.0% of the original's tokens per answer, interval 101.5 to 110.9. GSQ-RCO Q2_0 (ISTA-DASLab) at 2.40 bits produces 121.8%, interval 114.6 to 129.5."
  caption="Syzygy's evidence for long-horizon collapse: answer length on AIME relative to BF16, matched question by question. Its subtitle says 60 problems, while the card's AIME 2026 row says 30 × 8 (Syzygy, launch thread)."
/>

The chart is real evidence of something: a 2.40-bit GSQ-RCO build writes about 22% longer AIME answers than BF16 and Mach-2 about 6% longer, with intervals that don't overlap. Three things keep it from proving the thesis.

It measures length, not collapse. Nowhere does Syzygy publish an agentic score for GSQ-RCO Q2_0 or any other conventional quant, so the claim that they "collapse over long horizons" has no long-horizon measurement behind it. The comparison point is also an odd choice of "traditional": GSQ-RCO learns its grid assignment per weight and picks a type per tensor. ISTA-DASLab's own plot, as I read it [when we covered that build](/articles/big-moe-one-3090), puts Q2_0 at about 89 against the base's 93.12 on a short-benchmark average, so it isn't a model that kept its short-task scores intact either.

The subtitle says "Same 60 problems", and AIME 2026 is 30. Maybe the chart pools AIME 2025 and 2026; nothing says so.

And Mach-2's own interval, 101.5% to 110.9%, excludes 100. By the launch's own heuristic, Mach-2 also drifts, just less.

I think the thesis is probably right in direction. Compression error is not the same thing at step 1 and step 500, and evaluating a 1.7-bit model on 10-hour agent tasks is the right instinct. But "new Pareto frontier for in-RAM LLM compression on agentic workloads" needs the other builds on the same agentic harness, and that table doesn't exist yet.

## Next to the other Flash-Next builds

This model has become a common test subject, which makes Mach-2 easy to place. These numbers come from the respective sources as covered on this site, not from my own runs:

| build | bits on the backbone | backbone size | n-gram table | evidence |
|---|---:|---:|---|---|
| Mach-2 Additive Medium | 1.70 | 26.7 GB | BF16, 102.4 GB | 8 benchmarks vs BF16, incl. long agent runs |
| bnb2 ([Dettmers](/articles/runtime-dynamic-compression)) | 1.49 | 23.91 GB | not stated | WikiText-2 perplexity only |
| GSQ-RCO Coder ([ISTA-DASLab](/articles/runtime-dynamic-compression)) | 3.5, half the experts removed | about 29.6 GB resident | separate shard, can stay on disk | coding and SWE-bench retention |
| NVFP4 ([NVIDIA](/articles/qwen3-8-flash-next)) | 4 | not stated here | FP8, 47.7 GB | NVIDIA's own evaluation |

Two things stand out. Mach-2 is the only build here with long-horizon agentic results against its own BF16, which I'd weigh above its bit rate. And it is the most conservative treatment of the n-gram table anyone has shipped: NVIDIA's official release, which this site called the least compressed treatment of the table at the time, kept it at FP8; Syzygy kept it at BF16, twice that. If you want Mach-2 on a 64 GB machine, the table is the part to attack, and the Flash-Next article's survey shows 4- and 5-bit versions of it working elsewhere.

The same chart has a Mach build of DeepSeek V4.1 Flash at 2.0 bits and 146.8 GB, keeping 98.7% of the original's average. Those weights aren't released; Syzygy says they will serve it "through inference providers over the coming weeks", and replied to a question in the thread that they are "happy to open source it as well" if there's interest. Nothing about that one can be checked yet.

## What I'd do with it

If you have a 32 GB GPU and 128 GB of RAM, Mach-2 is the strongest-evidenced way I know to run Flash-Next's full 512-expert model at that budget, with the full n-gram table, and the weights are Apache-2.0. Build the fork with CUDA, run with `-ngl 99 -fa on --mlock`, and use the card's sampling settings (temperature 1.0, top_p 0.95, top_k 20); the card warns that without them "the model can loop". Check `GGML_MACH1_TIME=1` once to see the fused kernels engage.

If you're CPU-only or on AMD, wait: the codec runs on the CPU there, and the fork's own estimate is 3 to 5 times slower than Q4_K_M.

And if you're comparing quants, Syzygy's harness descriptions are the most useful part of the card. Their table makes the case that a sub-2-bit MoE can keep agentic scores within noise of BF16. It doesn't yet show that conventional quants fall apart where Mach-2 holds, and the one benchmark that clearly moved is the one the thesis says should have been safe.

## How I checked

- **Bits.** I read the headers of all 132 safetensors files in `SyzygyResearch/Mach-2-Additive-Medium` at commit `b0234bb` with HTTP range requests and summed parameters and bytes per group; the n-gram table's 22 shards were read the same way. Expert rates come from the `k4`–`k8` expert index tensors in each layer shard. I downloaded only `codebook.safetensors` (3 MB), `tlut.safetensors` and one 128-byte tile, and checked my reimplementation of the spine decoder against `decode.py`'s output on that tile.
- **GGUF.** I range-read the first 40 MB of `Mach-2-Additive-Medium.mach1.gguf` (commit `31b2cd6`), parsed its 61 metadata keys and 2,405 tensor descriptors, and summed tensor sizes by type.
- **Code.** I shallow-cloned `SyzygyResearch/llama.cpp-mach1` at `31261ba` (5 October 2026) and read `docs/mach1.md`, the Mach op contracts in `ggml/include/ggml.h`, the CPU reference in `ggml/src/ggml-cpu/ops.cpp`, `src/models/qwen4exp.cpp` and `mach1.cpp`, and the CUDA D4 kernels. I did not build or run it.
- **Papers.** QTIP's HYB code, Table 1 and figures are from arXiv 2406.11235v4. AQLM, QuIP# and VPTQ are cited for their mechanisms only.
- **Benchmarks.** Scores and footnotes are the card's. Task counts are score × tasks, which lands on whole numbers for six of the eight rows. The intervals are unpaired two-proportion intervals, so they overstate the noise for these paired runs and understate it for GPQA and AIME's repeated samples.
- **Not checked.** The 110 tokens/s on a 48 GB MacBook, every Qwen3.8-27B number beyond its plotted average, the DeepSeek build, how the per-expert rates were chosen, and whether the GSQ-RCO answer-length run used the 60 problems its subtitle names.
