~/satyajit

A 200 GB MoE on one RTX 3090: the experts live in RAM, the speed lives in the CPU

mdjsonmcp

2026-10-06 · 18 min · mixture-of-experts · inference-optimization · quantization · offloading · on-device · systems · vllm · qwen

Two posts went past on the same day. @0xSero wrote "Here's Deepseek-v4.1-Flash running on CPU + DDR4 with just a single RTX 3090", and in the next post listed what you need: one 3090, 256 GB of DDR4 and 500 GB of NVMe, with the repository 0xSero/dsv41-flash-offload. A few hours later @needmorevram posted a screen recording of "Qwen 3.8 Flash IQ3_S (GSQ RCO) on a single RTX 3090, this time paired with DDR5-5200 MT/s memory".

DeepSeek-V4.1-Flash carries 543.6 billion routed-expert parameters. The 3090 has 24 GB. Both runs work for the same reason, and the reason has a number attached. This piece derives that number from the model configs, checks it against what the repository measured, and then finds the place where the simple story stops being true.

0xSero/dsv41-flash-offload@0d99cdd · snapshot 2026-10-06
tracked files
109
license
MIT
branch
main
tests
1 file
source
309.9 kB
commit date
2026-10-05
source by language
Python184.2 kB(31)C48.9 kB(1)Shell42.0 kB(17)CUDA19.4 kB(2)C++12.4 kB(2)Dockerfile3.0 kB(1)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at 0d99cdd — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

A dark blue banner: 'DeepSeek-V4.1-Flash · local. One RTX 3090. Full DeepSeek. 200 GB MoE · experts across GPU, CPU and RAM · 262k context.' A card on the right reads 20 tok/s, 1 user (D141, default) and 41 tok/s, 4 users + Arc B70 (experimental). Footer: community build, not affiliated with DeepSeek.
The repository's own banner. 'Full DeepSeek' means the full model at 3.0 bits per weight, not the original FP8 and FP4 checkpoint; the 41 tok/s needs a second, Intel Arc B70 card and is marked experimental (0xSero/dsv41-flash-offload README).

What the repository actually runs

The task I set myself was to find the --override-tensor rules. There are none. This is not llama.cpp, ik_llama or KTransformers. The README pins its stack: vLLM 0.13 from a DeepSeek-V4.1 backport image, Ampere patches from a third party, the vllm_exl3 plugin with exllamav3 kernels built for sm_86, and a set of patches of its own that the image applies at build time and checks against SHA-256 receipts.

The weights are Mia-AiLab's EXL3 quant at an average 3.02 bits per weight, about 205 GB in 41 shards (reported, model card). EXL3 is exllamav3's format: a trellis code with a codebook (mul1), not the block-scaled integers of a GGUF Q4_K. That detail matters later.

Mia-AiLab/DeepSeek-V4.1-Flash-EXL3-3.0bpw@c5534b9 · snapshot 2026-10-06
parameters
109.55B
repo size
219.27 GB
architecture
DeepseekV41ForCausalLM
task
image-text-to-text
library
exllamav3
license
mit
safetensors
41 shards
largest file
8.12 GB
files
51
downloads
489
likes
5
parameters by dtype
BF161.18BF16471.3MF3242.4MI16107.85B
exl3exllamav3quantizedmoevllmdeepseekmultimodal

repo last modified 2026-09-12

The README's tier table says where everything lives (reported):

tierholdssize
VRAM, 24 GiB at 936 GB/sattention (with wo_a expanded to FP16), embeddings, shared experts, LM head, routers, hyper-connections, Engram projections~7.7 GiB
fp8 KV cache1.5 GiB (297,224 tokens)
CUDA graphs, prefill activations, DMA staging~5 GiB
mirror cache of the hottest expertswhat is left (8.76 GiB = 709 of 15,360 experts in run D107)
DDR4, pinned and GPU-mappedall 15,360 routed experts190.5 GiB
NVMe, memory-mappedEngram n-gram tables, fp8189 GiB

Three of those numbers I can check from the configs. DeepSeek-V4.1-Flash has 40 MoE layers, 384 routed experts per layer, hidden size 5,120 and expert width 2,304, and picks 6 experts per token (config.json). One expert is three matrices of 5,120 × 2,304, so 35,389,440 weights. Then (reasoned):

The Engram tables are DeepSeek's n-gram memory, the same idea as the 51B-parameter table in Qwen3.8-Flash-Next. They are read a few rows at a time (the README says about 50 small rows per token), so they stay on the NVMe and the page cache absorbs what it can. That is where the 500 GB of disk goes: 219.3 GB of EXL3 checkpoint plus 203.1 GB of the original model's last two shards, which hold the tables.

Why this works at all: decode reads only what it routes to

Generating one token at batch size 1 means reading every weight that token touches. On one stream the GPU has almost no arithmetic to do per byte, so the read sets the pace:

tok/s  ≤  Bbytes read per token\text{tok/s} \;\le\; \frac{B}{\text{bytes read per token}}

A dense model reads every weight. A routed MoE reads only the experts its router picked. For DeepSeek-V4.1-Flash that is 6 of 384 per layer, 1.6% of the routed weights. Per token, from the config (reasoned):

DeepSeek-V4.1-FlashQwen3.8-Flash-Next
routed experts, layers × top-k40 × 648 × 10
weights per expert35,389,4404,915,200
routed weights read per token8.49B2.36B
bytes per token at the run's quant3.19 GB at 3.0 bits1.03 GB at 3.5 bits (IQ3_S)
bytes per token at native FP4 (4.25 bits with scales)4.51 GBn/a

Qwen3.8-Flash-Next's shape comes from its own config.json: 48 layers, 512 experts of width 640 on a hidden size of 2,560, 10 picked per token. The 3.5 bits is the average the GSQ-RCO card gives for its IQ3_S file.

So the experts can sit anywhere that can deliver about 3 GB per token. Divide by the bus and you get the ceiling, if nothing else cost time (reasoned, peak theoretical bandwidths, DeepSeek at 3.0 bits):

where the routed experts are read frompeak GB/sDeepSeek ceilingQwen IQ3_S ceiling
DDR4-3200, 2 channels51.216 tok/s50 tok/s
DDR5-5200, 2 channels83.226 tok/s81 tok/s
DDR4-3200, 4 channels102.432 tok/s99 tok/s
DDR4-3200, 8 channels204.864 tok/s198 tok/s
PCIe 4.0 x16, GPU reading host memory~257.8 tok/s24 tok/s
RTX 3090 VRAM, if they fitted936294 tok/s907 tok/s

Two things fall out. First, a desktop's two DDR channels put DeepSeek at 16 to 26 tok/s before anything else is counted, which is why 0xSero's host is an AMD EPYC 7443P with eight. Second, streaming the experts over PCIe to the GPU is the worst option on the table: the link is a quarter of a desktop's RAM bandwidth. The fast path is the CPU computing the expert where it already sits.

The dense part does not go away. Attention, the shared expert, the routers and the LM head are read every token from VRAM. Taking the README's ~7.7 GiB of dense weights and removing the 1.32 GB embedding table (one row of it is read per token), the GPU reads about 6.9 GB per token, 7.4 ms at 936 GB/s (reasoned, an upper estimate). The widget below adds that to whichever side of the expert read is slower.

Decode ceiling with routed experts in system RAM · one RTX 3090 · reasoned
model
DDR4-3200, 2 channels · 51.2 GB/s8.2 tok/s · 122.1 ms/pass
DDR5-5200, 2 channels · 83.2 GB/s12.8 tok/s · 78.0 ms/pass
DDR4-3200, 4 channels · 102.4 GB/s15.4 tok/s · 64.8 ms/pass
DDR4-3200, 8 channels · 204.8 GB/s27.7 tok/s · 36.1 ms/pass
D107 measured 21.34 (black tick)
no CPU tier: misses over PCIe 4.012.1 tok/s · 82.6 ms/pass
One token reads 3.19 GB of routed experts at 3.00 bits per weight (DeepSeek-V4.1-Flash). Orange bars are bound by the RAM side, blue by VRAM. A ceiling assumes no two tokens in a pass share an expert and ignores compute, handshakes and the n-gram tables; measured numbers sit below it.

What sits on the GPU, and what does not

The split is the same in every engine that does this, whatever its flags are called:

In llama.cpp that split is usually spelled with --override-tensor (-ot) patterns that send the exps tensors to the CPU, or --n-cpu-moe. What llama.cpp's flags do not give you is the third bullet: a cache that follows the routing. That is the part 0xSero's repo adds, along with the CPU tier that serves the misses.

The cache: 4.6% of the experts, 41% of the reads

The VRAM mirror cache in ct/ct_vllm.py counts how often each (layer, expert) pair is routed to during decode, decays the scores, and every 16 steps swaps up to 48 of the coldest residents for the hottest non-residents. The home copy in pinned RAM stays valid, so a stale pointer can only be slower, never wrong. In run D107 the cache had 709 slots, 4.6% of 15,360 experts, and the README reports a hit rate of 0.408 (reported).

Routing is skewed enough that 4.6% of the experts serve 41% of the reads. With 6 experts routed per layer, 0.592 × 6 = 3.55 misses per layer (reasoned), which matches the README's "only about 3.5 experts per layer miss the VRAM cache".

A terminal coding-agent session with the model deepseek-v4.1-flash. The user asks the model to draw an ASCII diagram of an LLM; the reasoning trace and the beginning of a box diagram with Tokenizer, Embedding and Transformer x N layers are visible. The status line reads (omarchy-dsv41) deepseek-v4.1-flash and 5.1%/262k context.
The served model inside the Pi coding agent, from the post's own screen recording. The recording shows the model working; it shows no tokens-per-second figure, so the speeds in this article come from the repository's result files (0xSero's post, video frame).

The misses: a per-layer auction between PCIe and the CPU

A miss can be served two ways. The GPU can read the expert straight out of pinned host memory over PCIe ("zero-copy"), or the host CPU can compute it in place and send back only the small output vector. ft_split_k in ct/ft_tier_cu_v.cu decides, per layer, per step. It sorts the misses, tries every count kk of them to hand to the CPU, and keeps the kk that minimises the slower side:

t(k)=max⁡( thit (nh+nm−k)+tzc (nm−k),    ca+cb k )t(k) = \max\Big(\, t_\text{hit}\,(n_h + n_m - k) + t_\text{zc}\,(n_m - k),\;\; c_a + c_b\,k \,\Big)

where nhn_h and nmn_m are the layer's hits and misses. The constants are in the code, in milliseconds: thit=0.03t_\text{hit} = 0.03 per resident expert, tzc=0.58t_\text{zc} = 0.58 per zero-copy expert, ca=0.11c_a = 0.11 fixed per CPU job and cb=0.20c_b = 0.20 per CPU expert (the docker/entrypoint.sh default; ct_vllm.py falls back to 0.16, and D132 notes it was 0.24 before). This is the same PCIe-or-CPU trade the site worked through for FreeToken, and the source says so: the GPU side is a "vLLM port of kernels/cpu_avx2/ft_tier_cu.cu", and the campaign is called FreeToken-EXL3.

One MoE layer, one decode token · the split ft_split_k picks · constants from the repo
CPU ms per expert
0 to CPU · 4 over PCIe2.50 ms
GPU
CPU
1 to CPU · 3 over PCIe1.89 ms
GPU
CPU
2 to CPU · 2 over PCIe1.28 ms
GPU
CPU
3 to CPU · 1 over PCIe · chosen0.71 ms
GPU
CPU
4 to CPU · 0 over PCIe0.91 ms
GPU
CPU
Chosen: 3 of 4 misses on the CPU, 0.71 ms for this layer; over 40 layers that is 28.4 ms of routed-expert time per token, before attention and the dense weights. Each expert is 13.27 MB at 3.0 bits, so 0.58 ms over PCIe is about 23 GB/s and 0.20 ms on the CPU is about 66 GB/s.

Those constants are bandwidths in disguise. One expert is 13.27 MB at 3.0 bits, so (reasoned):

With four misses and two hits, the split sends three to the CPU and one over PCIe, and the layer costs 0.71 ms. Forty layers make 28 ms of expert time per token, and the dense read adds about 7.4 ms: a ceiling near 28 tok/s (reasoned). The README's own description is coarser, "the CPU computes them from DDR4 in about 1 ms per layer", which gives 40 ms and 25 tok/s.

Against what was measured

runwhat changeddecode, one streamsource
D030experts zero-copy over PCIe, no CPU tier, no cache4.10 tok/sresults/D030/sweep.json
D107+ AVX2 CPU tier + 709-slot VRAM cache21.34 tok/sresults/D107/sweep.json
D119262,144-token default, 8k to 261k prompts16.38-20.91 tok/sresults/D119/
D141current default19.3 tok/s (24.0 at 2 streams, 26.6 at 4)README

All reported; I did not run the image, which needs the hardware.

D030 is the PCIe row of the table above: a ceiling of 7.8 tok/s at 25 GB/s for the experts alone, plus the dense read, against 4.10 measured. D107 is the CPU tier: about 28 tok/s from the code's own constants against 21.34 measured. Both land at 55-76% of their ceilings, which is a normal place for a hand-built pipeline with a CPU/GPU handshake in every layer and NVMe reads in the middle.

Now the point I did not expect. If the CPU tier ran at the bus's peak, the 3.55 misses per layer would cost 1.89 GB per token, 9.2 ms on eight channels (reasoned). The code budgets 0.20 ms per CPU expert, three times that rate. The DDR4 bus is not the bottleneck on this machine. The CPU is.

The reason is the format. EXL3 stores each weight as a few bits of a trellis code that has to be decoded through a codebook before it can be multiplied. A Q4_0 block is a scale and sixteen bytes of nibbles; an EXL3 weight is arithmetic. On 22 Zen 3 cores with AVX2, that arithmetic runs out before the memory controller does. The Strata article found the same thing from the other side: its paper measures 23-26 GB/s on six cores for llama.cpp's i-quants, limited by codebook arithmetic, against about 42 GB/s for a plain 2-bit format with a hand-written kernel.

So the upgrade that helps 0xSero's build is more or faster cores, or more VRAM for the cache, not faster RAM. The repository's experimental branch is evidence for the second: an Intel Arc B70 added as a further expert tier takes four streams from 26.6 to 41.3 tok/s, with its quality A/B still pending (reported).

Prefill is a different machine

Decode touches 6 experts per layer per token. Prefill touches nearly all of them: the README says a 1,024-token chunk hits almost all 384 experts in all 40 layers, so the baseline moved about 203 GB over PCIe per chunk, about 8 s, and stayed near 95 tok/s. The fix is the opposite of decode's: make the chunk big (16,384 tokens in D061) and stream every expert to the GPU once per chunk with the copy engine, running busy experts as FP16 GEMMs. Prefill went to 690-790 tok/s from 8k to 261k (reported). A 261,000-token prompt still takes 341.6 s to its first token, and prefix caching turns a cached 131k-token agent turn into 0.76 s instead of 200.8 s.

During prefill the VRAM expert cache is released to make room and refilled after four decode steps. That is why the 8k row of D119 decodes at 16.38 tok/s and the longer rows near 20.5: the README says the cache is still re-warming after a short answer.

What it costs in quality

The banner says "Full DeepSeek". It is the full model at 3.02 bits per weight, and the repository is unusually direct about what that and its own kernels cost:

None of this is a comparison with the FP8 and FP4 original on a benchmark suite. Nobody has published one for this quant that I could find, so how much 3.02 bits costs on hard tasks is open.

The Qwen run on DDR5

The second post is only a video, but it is a screen recording of a dashboard, and the dashboard prints its numbers.

A terminal dashboard titled LLM VISUALS, process strata. Model Qwen3.8-Flash-Next-GSQ decoding at 99.0 tok/s; GPU RTX 3090 at 100%, 341 W, VRAM 23.3 of 24.0 GB; system RAM 60.3 of 67.2 GB. Context 10,608 of 262,144. Speculative draft-mtp: accept 87%, 3.29 per step. A request table lists seven requests; request 7 has a 2,421-token prompt and decodes at 97.0 tok/s, request 4 has a 142,550-token prompt and decoded at 54.1 tok/s.
A frame from the post's screen recording: the Strata engine serving the GSQ-RCO IQ3_S file on one RTX 3090 with 64 GB of DDR5-5200. The author says in the replies that the prefill column is not accurate (needmorevram's post, video frame).

What the frame shows (reported, read off the recording): the process is strata, the engine the site covered in TensorFold and Strata. VRAM is at 23.3 of 24.0 GB and system RAM at 60.3 of 67.2 GB. The MTP drafter is accepting 87% and committing 3.29 tokens per step. A 2,421-token prompt decodes at 97.0 tok/s; an earlier request with a 142,550-token prompt decoded at 54.1. In the replies the author says the CPU is an i5-13600K, and that "I tested the 3090 on both DDR4 and DDR5 and the difference was so tiny."

The file is ISTA-DASLab's GSQ-RCO IQ3_S: 3.50 bits on average over the transformer weights, a 54.8 GB first shard that must be resident and a 28.8 GB n-gram shard that can stay on disk (reported). GSQ learns the grid assignment per weight; RCO picks a quant type per tensor under a size budget. The card's plot puts IQ3_S at the base model's task average:

A line chart of task average (mean of AIME25, GPQA-Diamond and LiveCodeBench v6) against average bit-width. Q2_0 at 2.4 bpw scores about 89.1, IQ2_XS at 2.5 about 89.2, IQ3_XXS at 3.0 about 92.6, IQ3_S at 3.5 about 93.3, against a dashed base-model line at 93.12.
Quality against bits for the four GSQ-RCO files. IQ3_S sits on the base model's line; the drop is between 3.0 and 2.5 bits (ISTA-DASLab model card, task average vs bit-width plot).

Now the arithmetic (reasoned). At 97 tok/s and 3.29 tokens per step, one step takes 33.9 ms. A step that commits 3.29 tokens reads up to 3.29 × 1.03 = 3.4 GB of experts, if no two tokens share one. Two channels of DDR5-5200 move 83.2 GB/s at peak, so even with the CPU at full bus speed, at least 26% of those reads must come from the VRAM cache to fit in the step. If the CPU runs i-quants at Strata's own measured 23-26 GB/s, the cache has to serve about 78% of them. Strata's paper measured 72% on a 12 GB card; a 24 GB card holds more experts, so 78% is plausible. The OffloadCeiling widget's Qwen preset is that solution, not a measurement of the hit rate.

And "DDR4 and DDR5 … so tiny" is what the second case predicts. If the CPU tier runs at 25 GB/s, a bus that offers 51 or 83 GB/s is not the limit, and swapping it changes little. I did not measure the DDR4 run; the author did not post its number. One reply reports 93 tok/s on two 3090s and 128 GB of DDR4 (reported, unverified).

What to take from both

Related reading on the site: FreeToken for the PCIe-versus-CPU split as a measured policy, VRAM is a policy for what a residency choice costs when it is not stated, and GLM-5.3-Flash on four mining cards for another Ampere-only vLLM fork.

What I could not check

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "A 200 GB MoE on one RTX 3090: the experts live in RAM, the speed lives in the CPU", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026bigmoeone3090,
  author = {Satyajit Ghana},
  title  = {A 200 GB MoE on one RTX 3090: the experts live in RAM, the speed lives in the CPU},
  url    = {https://ai.thesatyajit.com/articles/big-moe-one-3090},
  year   = {2026}
}
share