2026-10-02 · 14 min · kv-cache · inference-optimization · long-context · attention · llm · open-source · explainer
A 1:26 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.
› transcript
Hi, I'm Oskar. A long chat fills a cache of past tokens, and it gets huge. Let me show you how to shrink it and still reuse it. The trick is to score the cache without looking at your question. Then one compressed cache can answer any question you ask later. First, read the whole context once and store it as a key value cache. Then ask the model to read that context again, straight from the cache. Give each stored pair a score: the most attention it gets during that reread. Keep the pairs that earned attention, and throw the rest away. The scoring never saw your question, so this one cache serves any query. The old way scores the cache against the current question, so it fits that one question and nothing else. Score by rereading instead, and the same small cache holds up across every later question. Here is the full cache, one cell per stored pair. Keep about the top thirty percent, and the answers barely change. That is three to four times less memory, and about twice the decode speed, with almost no loss. Score the cache without your question, and you compress it once and reuse it for every answer. So: reread to score the cache, compress once and reuse, and get it three to four times smaller and about twice as fast. Every source is in the full article. I'm Oskar. Bye!
This site has written about almost every lever on inference cost — quantization down to 1.5 bits, GPTQ in one fused loop, NVFP4 through Model Optimizer, linear and sparse attention, caching the value and rebuilding the key — and said almost nothing about the one structure that quietly decides what long context costs: the KV cache. This fills that gap.
The surfacing hook was a blog from RampLabs that took two KV-compression methods to a 320B sparse-attention frontier model and reported an 80% smaller cache. That is the applied data point at the end. The spine is the two papers underneath it, both with open code, both of which I read: KVzip (arXiv 2505.23416, NeurIPS'25 oral, MIT) and Fast KV Compaction via Attention Matching (arXiv 2602.16284, MIT). They attack the same cost from opposite ends.
What the KV cache is, and why it runs the bill
A decoder-only transformer caches, for every token it has already seen, the key and value vectors of every attention layer, so that generating the next token is one attention step over the cache instead of a re-read of the whole prompt. That is the trick that makes autoregressive decoding affordable. The cache is the receipt. For the mechanics of MHA, MQA, GQA and MLA byte by byte, the KV cache architecture note is the long version; the field guide to attention and how self-attention works sit under it.
The receipt grows linearly with context, and at long context it is the budget. The KVzip paper's own example: caching 120K tokens in Qwen2.5-14B at FP16 takes roughly 33 GB, more than the model's own 28 GB of parameters at the same precision (reported). And the cache is read in full on every decoded token, so it sets attention latency as much as attention memory. Shrinking it buys both.
The lazy way to shrink it is to drop tokens. The question every method answers differently is which tokens — and, the part that turns out to matter most, scored against what.
Query-aware eviction, and why a reused cache goes wrong
The established eviction methods — H2O, SnapKV, PyramidKV — score each KV pair by how much the current query attends to it, then keep the top fraction. That is query-aware scoring, and in a single-shot setting it works: the tokens the question looks at are the tokens the answer needs.
It breaks the moment you want to reuse the cache. Compress a long document's cache for the first question and you have kept the pairs that question attended to. Ask a second, unrelated question of the same document and those pairs are the wrong ones — the cache overfit the first query. The KVzip paper measures exactly this: SnapKV is strong when it re-runs prefill and compression per query, and falls off a cliff when its first-query cache is reused for later ones (reported, their Figure 2). The abstract's sharp version: query-aware methods "suffer from performance degradation even at a 90% cache budget ratio under multi-query scenarios" — a 10% eviction is enough to hurt (reported).
Query-agnostic eviction is the fix: score once, in a way that does not depend on any particular future query, so the compressed cache serves all of them. Then a long context is compressed a single time and reused across an entire conversation. The toy below is the whole argument. Compress the cache, pick which query you are asking, and watch what survives.
4 pairs Q2 needs were evicted (red outline).
Q2 · the warranty clause
In query-aware mode the cache is compressed once using Q1's attention, then reused. Q1 stays sharp; switch to Q2 or Q3 and watch its tokens turn red — they were evicted because the first query never looked at them. In query-agnostic mode the same budget keeps the backbone every query draws on, so all three hold up until the compression gets aggressive. That gap is the whole argument for compressing once and reusing.
Illustrative: synthetic 48-token context, three queries. The query-agnostic score models KVzip's reconstruction score as “what any query attends to”; the percentages are the toy's, not the paper's.
In query-aware mode the budget is spent on Q1; ask Q2 or Q3 and the tokens they need light up red, evicted because the first query never looked at them. In query-agnostic mode the same budget keeps the backbone every query draws on. The percentages are the toy's, not the paper's — but the shape is the result, and it is why "compress once, reuse" needs a query-independent score.
KVzip's score: make the model reread the context
So what is a good query-independent importance score? KVzip's answer is almost cheeky: ask the model to reconstruct the context from the cache, and keep whatever it attended to while doing so. The intuition is that a KV pair you cannot reproduce the context without is a pair some future query will probably need — and one that gets no attention during a full reconstruction is dead weight.
Concretely, after the normal prefill, KVzip runs one more forward pass over the cached context with a reconstruction instruction prepended. Its code states the prompt literally:
# model/wrapper.py — the reconstruction "query" is query-independent
prompt = "\n\nRepeat the previous context exactly."
# later chunks: "...starting with <last 8 tokens of the previous chunk>"
# the context tokens are then re-fed so the model attends back to the cache
input_ids.append((a_ids, torch.cat([q_ids, self.postfix_ids, a_ids], dim=1)))During that pass, each cached KV pair receives some attention from each reconstruction position. Its importance is the maximum attention it gets, across every query position and every grouped-query head that shares it. In the repository's scorer that is one line:
# attention/score.py — softmax over keys, then max over the GQA group and queries
attn_weights = nn.functional.softmax(attn_weights, dim=-1)
score = attn_weights[..., self.sink:self.sink + ctx_len].amax(dim=(-3, -2))Writing for the attention the -th reconstruction position places on cached pair after the softmax, the score is
Max, not sum: a pair that is decisive for even one reconstruction position is worth keeping. The sink tokens — the system prompt — are held out and never evicted, matching the paper. Then eviction is a global top-% over all layers and heads, with a non-uniform budget per head (KVzip adapts AdaKV's variable-length FlashAttention kernel for this), so a head that needs more pairs gets them.

Why reconstruction and not just the attention from the original prefill? Because prefill attention is spiky — it concentrates on a few tokens and under-scores many that the model still needs to reproduce the context later. The paper measures this gap directly (their Figure 5, prefill versus reconstruction attention). Score on the prefill pass and you evict pairs that matter; score on the reconstruction pass and you keep them.
Switch to prefill at a 40% budget: the signal is spiky, so the cache keeps a few heavily-attended tokens and throws away a pile the model actually needs to reproduce the context (the red bars). Reconstruction attention is broad because almost every token has to come back out, so the same budget keeps the pairs that matter. KVzip scores on the reconstruction pass for exactly this reason.
Illustrative: synthetic per-token scores, 48 tokens. The prefill-versus-reconstruction contrast is the paper's (Figure 5); the bar heights are the toy's.
The claims, checked against the paper and code
The headline numbers, each labelled:
- 3-4x smaller KV cache, ~2x faster FlashAttention decode, negligible loss (reported, abstract and their Figure 8). The efficiency figure is measured on LLaMA3.1-8B with a 124K-token context on an A100 in FP16, with the non-uniform head budget and variable-length FlashAttention-2. The 2x is the attention decode latency, not end-to-end.
- Accuracy holds while evicting up to 70% of the cache (reported). Their benchmark grid is the cleanest receipt the paper has: across twelve datasets, KVzip stays flat as the budget drops to ~0.3 while H2O, SnapKV and PyramidKV collapse — most visibly on retrieval, where a query-aware cache has thrown away the needle.

- Up to 170K-token contexts, LLaMA3.1 (3B/8B), Qwen2.5-7B-1M and 14B-1M, Gemma3-12B (reported). Models from 3B to 14B; GQA group sizes from 4 (LLaMA3.1-8B) to 7 (Qwen2.5-7B-1M); Gemma3's hybrid global/sliding-window attention is compressed only on the global layers, which dominate the cache at long context.
- The scoring is not free. The reconstruction pass costs about the same as one extra prefill — "approximately twice the computational overhead of standard prefill" — with under 2% additional memory, paid once per context (reported). Per chunk the scoring is and over the whole context , linear in context length with chunk size .
Two honest wrinkles the paper is up front about. First, the score needs a max over query positions after the softmax over keys, and that cross-dimensional dependency will not fuse into a block-wise FlashAttention kernel — so the scoring pass cannot use the same fused path decoding does, which is part of why it costs a second prefill. KVzip ships a softmax-free CUDA variant to claw some of that back, at a small quality cost (reported). Second, 3-4x is the context-dependent mode that pays the scoring overhead. KVzip also has a context-independent mode: compute head-level importance scores once per model and drop whole heads' context KV, DuoAttention-style, with no per-context overhead but a more modest ratio (0.6 recommended, so ~1.67x). That mode is also where KVzip quietly lands a second result — it replaces DuoAttention's head-score optimization, "tens of GPU hours," with a few forward passes inside a minute (reported) — two orders of magnitude at least (reasoned).
- license
- MIT
- branch
- main
- tests
- none found
- source
- 175.3 kB
- commit date
- 2026-02-11
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-02 at 5d84729 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow
shallow clone: counts describe the pinned tree, not the history
One number worth doing the arithmetic on. The paper stacks KVzip on QServe's 4-bit KV quantization: a 16-bit cache of 16.3 GB at 124K tokens becomes 1.2 GB with 4-bit quantization and 70% eviction together (reported) — about 13.6x off the 16-bit cache (reasoned). Eviction and quantization are orthogonal axes and they multiply.
Attention Matching: throw the pairs away and synthesize a smaller cache
KVzip keeps a subset of the real KV pairs. The companion method makes a different bet: keep none of them, and build a smaller set of synthetic keys and values that reproduce what the full cache would have done.
That is Attention Matching (Zweiger, Fu, Guo and Yoon Kim; code, MIT). Replace the full cache for one KV head with a shorter cache , , chosen so that the attention output is preserved even when the compacted prefix is concatenated with arbitrary later tokens — the user's next turn, the model's continuation. The trick is that attention over concatenated blocks is a mass-weighted mixture of each block's local attention, so it is enough to match two things per block, over a set of reference queries: the block's local attention output, and its attention mass.

The reason it is fast is that the objective decomposes into subproblems that have closed forms, so no gradient descent runs at compaction time. Construct the compacted keys first (by selecting from the real keys — the paper compares Highest-Attention-Keys and an Orthogonal-Matching-Pursuit selection), then solve a nonnegative least squares for per-key weights and set the bias — is "how many original keys' worth of attention mass this compact key carries" — then an ordinary least squares for the values . Reference queries come mostly from a repeat-prefill pass, with a few self-study prompts to broaden them, and layers are compacted one at a time so each later layer sees the queries the already-compacted earlier layers actually produce. The result: up to 50x compaction in seconds with little quality loss, two orders of magnitude faster than Cartridges (which trains a compact cache end-to-end over GPU-hours) at comparable ratios (reported, their Figure 1, QuALITY on Qwen3-4B). Stack it on a summary and it reaches ~200x total (reported). The paper runs KVzip as one of its one-shot eviction baselines, and Attention Matching generally comes out ahead on QuALITY and LongHealth — though it credits KVzip's non-uniform head budget with sometimes matching it at certain ratios (reported, their Figure 3). The two are rivals as much as companions.
The line worth flagging for what comes next: Attention Matching, as published, fits a separate set of keys and values for each KV head, and never addresses multi-head latent attention (MLA), which caches a shared low-rank latent instead of per-head keys and values. The two look complementary — compaction shrinks the sequence dimension, MLA shrinks the per-token dimension — but a method that picks a separate subset per head does not drop onto a shared-latent cache as written (reasoned; the paper does not discuss MLA).
What scaling to GLM-5.3 actually adds
Which is exactly where the applied data point gets interesting, because GLM-5.3-Flash is an MLA model, and a sparse one. Its eleven expensive layers are NoPE sparse MLA with a shared latent (kv_lora_rank: 512), and an indexer picks a top-2,048 subset of positions each query even looks at. Both compression methods assume things that architecture denies: KVzip's dense scoring reads all previous tokens, which sparse attention will not do; and both methods, in their original form, choose a separate subset of keys per layer-head pair, which a shared latent cache — one cache serving every head in the group — cannot honor.
The RampLabs write-up is their account of bridging that gap, and it is self-reported, with no code released, so I am reporting their claims as claims. They say they use dense scoring for KVzip with shared selection rules that balance importance across heads; and for Attention Matching they compare the two key-selection rules, then fit per-head biases and shared latent values against GLM's sparse attention, one layer at a time. The fit-per-head-biases-but-share-the-latent-values move is precisely the MLA extension the Attention Matching paper does not provide. Their headline: the KV cache 80% smaller while keeping over 90% of full-context accuracy on QuALITY (self-reported). Treat that as a vendor's own run on one benchmark — the useful signal is not the exact number but that a dense-attention, per-head method survived being bent onto a sparse, shared-latent frontier model at all.
Where it sits
Three axes, and they compose. Quantization makes each cached number smaller — 4-bit KV in QServe, the sub-2-bit expert work, GPTQ at speed. Architecture makes the cache grow slower — GLM-5.3's linear and sparse layers, SparDA prefetching its own KV, Grouped Value Attention, Flash-dLLM's cache for keys that move, Qwen3.8-Flash-Next's four changes. And compression after the fact throws cached pairs away: KVzip evicts the real ones it can prove are redundant; Attention Matching evicts all of them and synthesizes a shorter cache that reproduces the attention output. The two differ in what they keep — subset versus synthesis — and agree on the thing that made both worth writing up: the score is query-agnostic, so you pay to compress once and read the cache for free forever after.
The takeaway
The fix that makes KV-cache eviction actually reusable is not a cleverer notion of "important." It is refusing to let the current query define importance at all. KVzip gets its query-independent score by making the model reconstruct the context it already cached and watching what it reaches for; Attention Matching skips the question of which pairs to keep and fits a smaller cache to the attention output in closed form. Both are query-agnostic, both are open, and the frontier-model result that surfaced them is the least load-bearing part — a sign the ideas travel, bent once onto a 320B model by someone who had to serve it.