2026-10-06 · 27 min · quantization · mixture-of-experts · qwen · llama-cpp · gguf · benchmarks
Why read this
Notabletop 60%Decodes Mach-2's packed format: a QTIP-style trellis, not additive codebooks; 1.70 bits only without a 102 GB BF16 table; one benchmark gap above noise.
- Analysis found nowhere else
- Concrete numbers to act on
- Open code or weights
Quantization & compressionNeeds a workstation GPUApache-2.0Practitioner model
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 2 of 3: Open code or weights with real limits
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 3 of 3: The only place this analysis exists
Score 62 of 100, ranked 203 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
This site has now covered five different ways of shrinking Qwen3.8-Flash-Next, and every one of them had to decide what to do with the model's 51B-parameter n-gram embedding table. So when Syzygy announced a build "at 1.7 bits per weight on RAM", "9.4x smaller than its base model", that beats Qwen3.8-27B on DeepSWE, NL2Repo, Terminal-Bench and AutomationBench, the first thing I did was add up the files.
They come to 129 GB.
That isn't a contradiction, and the model card is upfront about why. But the number is a good way in. Mach-2 Additive Medium is a careful piece of engineering wrapped in two pieces of framing. "Additive" names a codec family it isn't really a member of, and "1.7 bits on RAM" describes the part of the model that sits in VRAM. Under the framing is a QTIP-style trellis quantizer with per-expert bit rates, which I think is the most interesting thing Syzygy shipped and the thing the launch never mentions.
- architecture
- Qwen4ExpForConditionalGeneration
- license
- apache-2.0
- safetensors
- 132 shards
- largest file
- 4.80 GB
- files
- 147
- downloads
- 43
- likes
- 1
The 26.7 GB is the text weights only. The repo also carries the n-gram embedding table at BF16 (22 shards, 102.4 GB) and a 1.5 GB 4-bit MTP draft for MLX. The GGUF repo is one 130.1 GB file holding both.
repo last modified 2026-10-06
129 GB, of which 26.7 are the 1.7 bits
The native repo is a set of packed safetensors files plus a decode.py that turns them back into a BF16 checkpoint. I read the header of every one of the 132 .safetensors files with HTTP range requests and summed what they hold. The text model is 125,743,653,795 parameters in 26,717,083,816 bytes of tensors, which is 1.6998 bits per weight. MANIFEST.json says 125,743,653,760 and 26,717,634,752 bytes, the difference being file headers and a few int64 offsets. So the headline bit rate is real, to four digits.
Then there is packed/table/: 22 files, 128 BF16 tensors of shape [2500012, 160], holding 51,200,245,760 parameters in 102,400,491,520 bytes. It is the n-gram embedding table from the base model, the trigram lookup the Flash-Next technical report spends a whole section justifying. Syzygy did not compress it at all. The card says so in its first sentence, "plus the model's n-gram embedding table at bf16 (102.4 GB)", and the GGUF's README tells you to --mlock it into about 105 GB of free host RAM.
So the 9.4x needs its denominator stated. BF16 text weights are 125.74B × 2 bytes = 251.5 GB, and 251.5 / 26.7 = 9.41. Correct. Put the table on both sides, since the BF16 base ships it too, and the ratio is 353.9 GB against 129.1 GB, or 2.74x. Averaged over all 176.9B parameters the files hold, the bit rate is 5.84.
Which is the right number depends on what you're asking. The table is read 16 rows per token, 5,120 bytes, so for decode bandwidth it barely exists. For "will it fit on my machine", it is four fifths of the download. "1.7 bits per weight on RAM" is the odd one: following Syzygy's own instructions, the 1.7-bit part goes to the GPU and the BF16 part is what sits in RAM.
- Routed experts
- 23.64 GB · 88.5% · trellis codes, 1.0 to 2.0 bits each, mean 1.57
- Attention, linear attention, shared experts
- 1.46 GB · 5.5% · trellis codes, 4.0 bits
- Head, embedding, mixers, router, norms
- 1.62 GB · 6.1% · int4, int5, int8 and BF16
Stored, the experts are 88% of the 26.7 GB. Everything that is not an expert costs 4 bits or more per weight.
The ledger also shows where the bits actually go. The routed experts are 120.8B of the 125.7B text parameters, 96.1%, and they average 1.565 bits. Everything else costs four bits or more. The attention, linear-attention and shared-expert matrices (2.91B parameters) are at 4.005 bits, the LM head at 5.25 (int5, group 64), the token embedding at 4.5 (int4 with an FP16 min and max per 64), the hyper-connection mixers at about 8, and 78.4M router, norm and recurrent-scalar parameters stay at BF16. That's the usual shape of an extreme quant: sub-2-bit numbers are only possible on a mixture of experts, because there the experts are almost the whole model.
The experts don't all get the same rate, either. Each expert's three matrices carry a rate_k4 byte, bits per group of four weights, and the shards bucket experts by it. Across all 48 layers and 24,576 experts:
| bits per weight | 1.0 | 1.25 | 1.5 | 1.75 | 2.0 |
|---|---|---|---|---|---|
| experts | 3,628 | 6,530 | 2,774 | 3,278 | 8,366 |
| share | 14.8% | 26.6% | 11.3% | 13.3% | 34.0% |
A third of the experts get two bits and one in seven gets one. The allocation also leans by depth: layer 0's experts average 1.69 bits, and the last nine layers sit between 1.48 and 1.53. How Syzygy chose the rates isn't documented. The exporter isn't public, and the fork's own docs say so: "The exporter that produces the code streams is not part of this repo."
A trellis wearing an additive name
"Additive" has a specific meaning in quantization, and I expected to find it. AQLM (Egiazarian et al., "Extreme Compression of Large Language Models via Additive Quantization") stores each group of eight weights as the sum of codewords drawn from several learned codebooks. Two codebooks of 256 entries, say, give 65,536 representable vectors for 16 bits of index. This is multi-codebook additive quantization. It sits in a line of vector quantizers next to QuIP#, which rotates weights with a randomized Hadamard transform until they look Gaussian and then rounds groups of eight to a fixed E8 lattice codebook, and VPTQ, which learns vector codebooks with second-order information.
Mach-2 doesn't do this. No tensor in either repo is a stack of codebooks whose entries get summed. What the files hold is the next step in that same line, trellis-coded quantization, and its closest published relative is QTIP ("Quantization with Trellises and Incoherence Processing", Tseng, Sun, Hou and De Sa).
From one code per weight to one path per tile
The reason anyone bothers with any of this is the distortion you pay per bit. Most GGUF quants are scalar: each weight gets its own small integer and a group shares a scale. (The IQ2 and IQ3 types are the exception; they borrow lattice codebooks from QuIP#.) At two bits, a scalar quantizer has four levels per weight and no way to exploit the fact that weights come in groups. QTIP's Table 1 puts numbers on it for an i.i.d. Gaussian source at 2 bits: the optimal scalar (Lloyd-Max) quantizer reaches 0.118 mean squared error, QuIP#'s 8-dimensional E8P codebook 0.089, a 256-dimensional trellis code 0.069, against an information-theoretic floor of 0.063.
Vector quantization gets its gain by quantizing eight weights jointly, but a codebook for 8 dimensions at 2 bits has 2^16 entries, and going to 16 dimensions would need 2^32. A trellis gets most of the high-dimensional gain without that table. Instead of choosing a codeword per group, it chooses a path through a state machine for a whole sequence of weights, and each state emits a value. Viterbi search finds the path with the least squared error; storage is just the bits that steer it.

QTIP's practical contribution is the "bitshift" trellis. If each state is just the last L bits of the stream, the next state is the old one shifted left by k bits with k fresh bits on the end. Any state can be read straight off the bitstream at position t × k, so decoding is parallel and the trellis graph never has to be stored. The remaining cost is the state-to-value table, 2^L entries for L = 16, and QTIP's second contribution was to compute values from the state with a hash rather than look them all up.


What the spine stores
The non-expert matrices are where the lineage is easiest to see, because they use QTIP's hybrid code almost verbatim. Each 16 × 16 tile is 64 int16 words, 1,024 bits for 256 weights, so 4 bits per weight. decode.py builds the value table like this:
NE_K, NE_L, NE_V, NE_TLUT_BITS, NE_TD = 4, 16, 2, 9, 16
def _ne_full_lut(tlut):
...
s = np.arange(1 << NE_L, dtype=np.int64)
p = s * (s + 1)
row = (p >> (16 - NE_TLUT_BITS - 1)) & ((1 << NE_TLUT_BITS) - 1)
table = small[row].copy()
table[:, 0] *= (1 - ((p >> 15) & 1) * 2).astype(np.float32)Compare QTIP's Algorithm 3, "HYB": hash the state as x · x + x, use bits from position 15 − Q upward as an index into a 2^Q × 2 codebook, and flip a sign from bit 15. With Q = 9, that's the same code, down to which bits it takes. The 512-row table is packed/ne/tlut.safetensors, and when I loaded it every entry turned out to be an integer between −32 and 31 times one step of 0.0867, which is what the fork's ggml.h means when it calls this the "rotated int-lattice trellis". Each 16-bit window yields two weights, so the window advances 8 bits a step, and the last windows wrap around to the start of the tile. QTIP calls that tail-biting; it saves storing a separate start state.
The widget below runs this decoder on the first real tile of layer 0's shared-expert gate projection, using the shipped words and the shipped table. Its first two outputs, −1.560 and −0.693, are what decode.py produces for the same bytes before the tile's scale and rotations are undone.
window = bits 0 to 15 of 1024 · bold = the 8 fresh bits; 8 are shared with the previous step
- state s
- 0000011010100011 = 1699
- hash p = s(s+1)
- 2888300
- row = bits 6..14
- 73 → (-18, -8) × 0.0867
- bit 15 flips
- no
- weights 0, 1
- -1.560, -0.693
decoded tile, blue +, orange −, before scale and rotation
Slide it and watch the window. Half of every state is the previous state's bits, which is the trellis constraint at work: consecutive pairs of weights are not free to be anything, they're linked through the shared bits, and the encoder's Viterbi search exploits that link to get below what independent rounding could do at the same rate.
What the experts store
The routed experts use a different codec, which the GGUF metadata names: {"codec": "d4", "L": 16, "V": 4, "rates_k4": [4, 5, 6, 7, 8]}. Same bitshift idea, a 16-bit window over a tail-biting stream, but each state now emits four weights and the window advances K4 bits, where K4 is the expert's rate_k4. At K4 = 4 that's one bit per weight, at K4 = 8 two bits. decode.py reads it directly:
def _ex_states(words, K4):
words = np.asarray(words).view(np.uint16).astype(np.int64)
rows, nstep, step = words.shape[0], EX_TD * EX_TD // EX_V, int(K4)
bits = ((words[:, :, None] >> np.arange(15, -1, -1)) & 1).reshape(rows, -1)[:, :step * nstep]
bits = np.concatenate([bits, bits[:, :EX_L - step]], axis=1)
w = 1 << np.arange(EX_L - 1, -1, -1, dtype=np.int64)
idx = np.arange(nstep)[:, None] * step + np.arange(EX_L)[None, :]
return bits[:, idx] @ wThe state indexes a 65,536 × 4 table instead of a hash, shipped in packed/experts/codebook.safetensors (one table for 1 bit, one for 2 bits, a shared one for the rates between). I downloaded the 3 MB file to look at the entries. They are also small integers times a step: for the 1.0 to 1.75-bit rates, every coordinate is 0, ±1, ±2 or ±4 steps, and only about 2,400 distinct four-vectors appear among the 65,536 states; the 2-bit table adds ±3 and ±6 and has 14,571 distinct vectors. "d4" is not the D4 lattice, by the way. D4 points have an even coordinate sum, and barely half of these do. I read the name as "four-dimensional".
Varying the rate per expert is the part I'd call new. QTIP's published models use one rate per model, and Tim Dettmers' runtime dynamic compression paper explicitly blamed QTIP's weak showing below 2 bits on fractional rates being impractical, since encoding a matrix "can take days of GPU time". Syzygy's answer is five rungs, each a different K4 on the same 16-bit trellis, assigned per expert. Whatever their exporter does to pick the rung, it ran over 24,576 experts.
Rotations, signs from a hash, and tiles that share a scale
Trellis codes assume Gaussian-looking inputs, and real weight matrices have outliers. Mach-2 handles that the way QuIP# and QTIP do: each matrix is multiplied on both sides by a random-sign diagonal and a Hadamard matrix before quantizing, and decoding undoes it, W = diag(sv) · H · Ŵ · H · diag(su) in the fork's notation. Two details caught my eye.
The dimensions here aren't powers of two. An expert's matrices are 640 × 2560, and 2560 = 20 × 128 while 640 = 20 × 32. decode.py builds 12- and 20-point Hadamard matrices from quadratic residues (Paley's construction) and takes Kronecker products with ordinary power-of-two butterflies. The CUDA kernels hard-code 12 × 12 and 20 × 20 sign tables for the same job.
And in the native format, the expert sign vectors aren't stored. sign_seed() hashes a string like canon|L12|e301|gate|SU with SHA-256, and expert_signs() feeds the first seven bytes into what the constants identify as a Philox4x32-10 counter generator. 24,576 experts × 3 matrices × 2 sides of signs are regenerated at load. The GGUF stores them instead, as FP16, which costs 0.47 GB.
The last piece is wave_gamma: per expert and matrix, 200 FP16 scales, one per anti-diagonal of the 40 × 160 grid of tiles (199 diagonals; one of the 200 slots goes unused). Each tile's decoded values are multiplied by its diagonal's scale before the rotations. I don't know why anti-diagonals; it is cheap side information, 400 bytes per matrix against 410 KB of codes at 2 bits.
So where is the "additive"?
Nowhere in the Mach-2 math. In the fork, "additive" is a label on an older payload: src/models/mach1.cpp describes format version 3 as "additive (V8 experts, rotated NE spine, int5-g64 head, nibble-LUT embed)", and version 2 added a shared low-rank correction to "demoted" experts, y = B (c_e ⊙ (A x)), which is genuinely additive in the residual sense. Mach-2 is format version 5 and uses neither. My best guess is that the name came along with the product line. A second possibility is Syzygy's own site, which advertises "Multiplication-Free Models at 1.7 Bits per Weight". Codebook coordinates of 0, ±1, ±2 and ±4 can be applied with shifts and adds. But that's my reading, not theirs, and the shipped kernels don't do it: the CPU path multiplies with int8 dot-product instructions, and the CUDA path with FP16 tensor cores. Mach-2 also has a 4-bit spine and an int5 head that are ordinary multiplies however you build them.
Decoding a matrix you never materialise
A packed checkpoint is only useful if you can run it without first expanding it back to BF16, and the llama.cpp-mach1 fork is where that happens. The codec is not a ggml quant type. The fork adds custom ops (GGML_OP_MACH1_D4_MM for experts, GGML_OP_MACH1_RT_MM for the spine, GGML_OP_MACH1_HEAD_MM for the head) and stores the code streams under plain int16 and int8 tensor types. Stock llama.cpp can't load the file, and the docs warn that the fork's ggml libraries must never be mixed with stock builds, because "the op enum differs - an ABI break, not a graceful failure".
- license
- MIT
- branch
- master
- tests
- 214 files
- source
- 33.1 MB
- commit date
- 2026-10-05
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-07 at 31261ba — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
The CPU reference makes the trick clear. For each (token, expert) pair, ggml_compute_forward_mach1_d4_mm (ggml/src/ggml-cpu/ops.cpp) multiplies the input by the expert's su signs and runs a Hadamard transform on it, so the input lives in the same rotated space as the codes. It then quantizes that rotated vector to int8 in groups of 16, walks each tile's trellis to get 64 four-vectors of 4-bit integers, multiplies the integer codes against the int8 activations with dpbusd-style instructions, and finally scales, rotates back and applies sv. The decoded matrix never exists in memory. The inner decode is three lines:
for (int j = 0; j < 64; ++j) {
const int b = j*K4;
const int by = b >> 3;
const uint32_t w24 = ((uint32_t) buf[by] << 16) | ((uint32_t) buf[by + 1] << 8) | buf[by + 2];
z[j] = z8[(w24 >> (8 - (b & 7))) & 0xFFFFu];
}Read 24 bits at the step's bit offset, shift out the 16-bit state, look up four packed nibbles. On CUDA, the walk kernel decodes tiles into shared memory as FP16, split into a high and a low half to keep precision, and feeds them to tensor cores with wmma. The fork also pattern-matches whole regions of the decode graph into fused kernels (gdn_full, shexp_gu, qkvzb, exp_mega), and its docs tell you to check a kernel census with GGML_MACH1_TIME=1, because if those fusions don't engage, "decode will be much slower than it should be".
The honest line in the docs is about the CPU: "usable but slow relative to llama.cpp's hand-tuned quant kernels - roughly 3x behind Q4_K_M decode on AVX2 hosts and ~5x on AVX-512". A trellis walk costs integer work per weight that a Q4_K block doesn't, and on a CPU that is enough to lose despite reading far fewer bytes. The fork's verdict is "A GPU is strongly recommended", and for CUDA it's the only fast path; AMD builds with HIP "run the codec ops on the CPU". The fork's Vulkan shaders cover the Mach ops but aren't certified bit-identical to the CPU reference.
Memory and speed, from the bytes
What follows is for the fork and the GGUF.
The GGUF holds more than the card's 26.7 GB. I parsed its header (2,405 tensors) and summed every tensor: 130.07 GB, of which the table is 102.40 and the rest 27.67 GB. The extra gigabyte over the native format is mostly the stored sign vectors, the embedding in its lossless nibble-LUT form, and the router kept at F32. With -ngl 99 those 27.67 GB go to the GPU, so the card's "32 GB or more" is right with a few gigabytes to spare.
The KV cache is small, because only 12 of the 48 layers are full attention, each with 2 KV heads of 256: 24,576 bytes per token at FP16, plus a little for the sparse-attention indexer's keys. That's 0.81 GB at a 32,768-token context and 6.44 GB at the full 262,144. The 36 Gated DeltaNet layers add a fixed recurrent state of about 113 MB at FP32. Long agent runs are where this model's architecture already helps, and the compression leaves that alone.
The table goes to host memory. The fork tags it as a layer input, the same placement llama.cpp gives token embeddings, so with --mlock you need about 105 GB of free RAM and ulimit -l unlimited. Without --mlock it is memory-mapped and pages the OS drops are read back from disk. At 16 rows of 320 bytes per token the bandwidth is nothing; the cost is latency on cold pages. So the practical box is a 32 GB GPU plus 128 GB of RAM and 130 GB of disk. Syzygy says they checked it on one A100 80 GB.
Speed is where I'd temper expectations, and the reason is in the third ledger view. Decoding one token touches 10 of 512 experts in each layer but every non-expert matrix, the head and the mixers. By my count that's 3.18 GB per token, and the experts are only 0.46 GB of it, 14.5%. The 4-bit spine alone is 1.46 GB. For comparison, the same touched parameters at a flat 4.5 bits would be 3.75 GB. So against a conventional 4-bit build, Mach-2 cuts the bytes per token by about 15%, while cutting the model's resident size by more than half. It is a capacity codec far more than a bandwidth codec, and the trellis decode work comes on top.
Dividing memory bandwidth by 3.18 GB gives a ceiling, not a prediction: about 86 tokens/s on an M4 Pro's 273 GB/s, 172 on an M4 Max's 546 GB/s, and 641 on an A100's 2,039 GB/s, with one token per step. The launch thread claims "over 110 tokens per second on a 48GB MacBook". That comes from Syzygy's own macOS engine, not the fork, and nothing names the chip, context or settings. On an M4 Pro it is above the single-token ceiling, which would need speculative decoding; the repo does ship a 4-bit copy of the base model's MTP layer "for speculative decoding with MLX". On an M4 Max it fits without. A 48 GB Mac also cannot hold the 102 GB table, so that run must page the table from SSD. I couldn't check the 110.
Eight benchmarks, one real gap
The card's table compares Mach-2 with Qwen3.8-Flash-Next at BF16, "both run on the same harness, settings and tasks", and the footnotes are unusually specific: Terminal-Bench 2.1 with Terminus 2 and four hours per task; DeepSWE with Claude Code as the agent, up to 10.5 hours and 700 steps per task; NL2Repo in OpenHands without internet; τ³-Banking with GPT-5.4 mini as simulated user and judge. Running a 1.7-bit model through 10-hour agent episodes is expensive, and most quantization releases stop at perplexity. Credit for that.

The averages check out. Mach-2's four scores average 58.215 and BF16's 60.43, a ratio of 96.3%, as plotted. The more useful thing I could do with the table was turn percentages back into task counts, because most of them land on whole numbers. Humanity's Last Exam is 731 against 860 correct out of 2,158. Terminal-Bench is 70 against 73 out of 89. DeepSWE is 55 against 56 out of 113. AIME is 231 against 230 answers out of 240.
Grey is a rough 95% band for the difference. Only Humanity's Last Exam, the knowledge exam, falls outside it. Every agentic gap is a few tasks wide and sits well inside its noise.
Put a rough 95% interval around each difference and only one gap survives. HLE's 5.98-point drop is about four standard errors. Every agentic difference is a handful of tasks inside an interval of 5 to 14 points: a model that matched BF16 exactly would produce gaps this size. The interval is crude (I can't pair the runs, because per-task results aren't published, and pairing would tighten it), but it says the agentic rows are consistent with no loss, which is a weaker claim than the 98.2% DeepSWE retention suggests. AutomationBench at 102.0% and AIME at 100.4% are the same noise pointing the other way.
The one difference that is real sits on the knowledge exam. HLE is a single-turn test of facts and reasoning, exactly the kind of "short benchmark" the launch says traditional quants preserve. A plausible reading is that squeezing 96% of the parameters to about 1.5 bits costs stored knowledge first and agentic behaviour later, if at all. That fits what other heavily compressed models have shown, but this table is the evidence I have, and it is one benchmark.
The comparison with Qwen3.8-27B is harder to check. The post says Mach-2 outperforms it "on benchmarks such as DeepSWE v1.1, NL2Repo, Terminal Bench, and Automation Bench". What Syzygy published for the 27B is a single point on the chart: 51.5 averaged over HLE, Terminal-Bench, AutomationBench and DeepSWE. NL2Repo isn't in that average, and no per-benchmark 27B scores appear in the post, the thread or either card. A 6.7-point lead on a four-benchmark average is believable for a 125B-parameter MoE against a 27B dense model. The per-benchmark claim I couldn't verify.
The long-horizon argument, and what backs it
The launch's thesis is the interesting sentence: "Traditional quantizations fall short. They preserve short benchmark performance but collapse over long horizons." It's a plausible mechanism. Small per-token errors compound over a long trajectory, a model that drifts loops or overthinks, and a 50-step agent run gives drift more chances than a single answer does.
The evidence offered is one chart, posted in the thread with the line "A strong heuristic for long horizon performance retention is response length, typically increased by overthinking."

The chart is real evidence of something: a 2.40-bit GSQ-RCO build writes about 22% longer AIME answers than BF16 and Mach-2 about 6% longer, with intervals that don't overlap. Three things keep it from proving the thesis.
It measures length, not collapse. Nowhere does Syzygy publish an agentic score for GSQ-RCO Q2_0 or any other conventional quant, so the claim that they "collapse over long horizons" has no long-horizon measurement behind it. The comparison point is also an odd choice of "traditional": GSQ-RCO learns its grid assignment per weight and picks a type per tensor. ISTA-DASLab's own plot, as I read it when we covered that build, puts Q2_0 at about 89 against the base's 93.12 on a short-benchmark average, so it isn't a model that kept its short-task scores intact either.
The subtitle says "Same 60 problems", and AIME 2026 is 30. Maybe the chart pools AIME 2025 and 2026; nothing says so.
And Mach-2's own interval, 101.5% to 110.9%, excludes 100. By the launch's own heuristic, Mach-2 also drifts, just less.
I think the thesis is probably right in direction. Compression error is not the same thing at step 1 and step 500, and evaluating a 1.7-bit model on 10-hour agent tasks is the right instinct. But "new Pareto frontier for in-RAM LLM compression on agentic workloads" needs the other builds on the same agentic harness, and that table doesn't exist yet.
Next to the other Flash-Next builds
This model has become a common test subject, which makes Mach-2 easy to place. These numbers come from the respective sources as covered on this site, not from my own runs:
| build | bits on the backbone | backbone size | n-gram table | evidence |
|---|---|---|---|---|
| Mach-2 Additive Medium | 1.70 | 26.7 GB | BF16, 102.4 GB | 8 benchmarks vs BF16, incl. long agent runs |
| bnb2 (Dettmers) | 1.49 | 23.91 GB | not stated | WikiText-2 perplexity only |
| GSQ-RCO Coder (ISTA-DASLab) | 3.5, half the experts removed | about 29.6 GB resident | separate shard, can stay on disk | coding and SWE-bench retention |
| NVFP4 (NVIDIA) | 4 | not stated here | FP8, 47.7 GB | NVIDIA's own evaluation |
Two things stand out. Mach-2 is the only build here with long-horizon agentic results against its own BF16, which I'd weigh above its bit rate. And it is the most conservative treatment of the n-gram table anyone has shipped: NVIDIA's official release, which this site called the least compressed treatment of the table at the time, kept it at FP8; Syzygy kept it at BF16, twice that. If you want Mach-2 on a 64 GB machine, the table is the part to attack, and the Flash-Next article's survey shows 4- and 5-bit versions of it working elsewhere.
The same chart has a Mach build of DeepSeek V4.1 Flash at 2.0 bits and 146.8 GB, keeping 98.7% of the original's average. Those weights aren't released; Syzygy says they will serve it "through inference providers over the coming weeks", and replied to a question in the thread that they are "happy to open source it as well" if there's interest. Nothing about that one can be checked yet.
What I'd do with it
If you have a 32 GB GPU and 128 GB of RAM, Mach-2 is the strongest-evidenced way I know to run Flash-Next's full 512-expert model at that budget, with the full n-gram table, and the weights are Apache-2.0. Build the fork with CUDA, run with -ngl 99 -fa on --mlock, and use the card's sampling settings (temperature 1.0, top_p 0.95, top_k 20); the card warns that without them "the model can loop". Check GGML_MACH1_TIME=1 once to see the fused kernels engage.
If you're CPU-only or on AMD, wait: the codec runs on the CPU there, and the fork's own estimate is 3 to 5 times slower than Q4_K_M.
And if you're comparing quants, Syzygy's harness descriptions are the most useful part of the card. Their table makes the case that a sub-2-bit MoE can keep agentic scores within noise of BF16. It doesn't yet show that conventional quants fall apart where Mach-2 holds, and the one benchmark that clearly moved is the one the thesis says should have been safe.
How I checked
- Bits. I read the headers of all 132 safetensors files in
SyzygyResearch/Mach-2-Additive-Mediumat commitb0234bbwith HTTP range requests and summed parameters and bytes per group; the n-gram table's 22 shards were read the same way. Expert rates come from thek4–k8expert index tensors in each layer shard. I downloaded onlycodebook.safetensors(3 MB),tlut.safetensorsand one 128-byte tile, and checked my reimplementation of the spine decoder againstdecode.py's output on that tile. - GGUF. I range-read the first 40 MB of
Mach-2-Additive-Medium.mach1.gguf(commit31b2cd6), parsed its 61 metadata keys and 2,405 tensor descriptors, and summed tensor sizes by type. - Code. I shallow-cloned
SyzygyResearch/llama.cpp-mach1at31261ba(5 October 2026) and readdocs/mach1.md, the Mach op contracts inggml/include/ggml.h, the CPU reference inggml/src/ggml-cpu/ops.cpp,src/models/qwen4exp.cppandmach1.cpp, and the CUDA D4 kernels. I did not build or run it. - Papers. QTIP's HYB code, Table 1 and figures are from arXiv 2406.11235v4. AQLM, QuIP# and VPTQ are cited for their mechanisms only.
- Benchmarks. Scores and footnotes are the card's. Task counts are score × tasks, which lands on whole numbers for six of the eight rows. The intervals are unpaired two-proportion intervals, so they overstate the noise for these paired runs and understate it for GPQA and AIME's repeated samples.
- Not checked. The 110 tokens/s on a 48 GB MacBook, every Qwen3.8-27B number beyond its plotted average, the DeepSeek build, how the per-expert rates were chosen, and whether the GSQ-RCO answer-length run used the 60 problems its subtitle names.