~/satyajit

Context Language Models: the bitter lesson comes for context management

mdjsonmcp

2026-10-02 · 16 min · explainer · agents · llm · long-context · inference-optimization · kv-cache · reinforcement-learning

A 1:42 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.

› transcript

Hi, I'm Vesper! Let me show you how a model can run its own memory, by treating its context as a file it rewrites. Instead of a harness deciding what to forget, the model edits its context like a file: append, rewrite a line in place, or offload a chunk. A turn starts with the live context. It is mirrored as a plain text file the model can open. The model reads the file and edits it with Bash. It can append, rewrite a line in place, or move a chunk out to storage. The edit syncs back and becomes the context for the next turn. A fixed harness summarizes on a schedule. One lossy gist, and the exact fact you needed is gone. The model instead keeps the right lines and drops the rest, in any strategy it can write as code. With no training, this beats the strongest harness on a deep research benchmark, and gets there using less compute. Online reinforcement learning folds the strategy into the weights. A small model jumps a lot, and spends less compute doing it. An arbitrary edit is not free to serve. Standard serving re-prefills the whole tail after the edit, even text that never changed. Suffix Cache Reuse keeps the unchanged tail's cache and splices it back, so only the edited part is recomputed. Context management was the last piece of the agent loop written by hand. The model can learn to do it better. Context is a file. The model reads it, edits it, and keeps what matters. That's the whole idea. Thanks for watching! Every source is in the full article. I'm Vesper. Bye!

Every long-running agent eventually runs out of room. The transcript grows, the context window fills, and something has to decide what to throw away. Today that something is almost never the model. It is the harness — the loop wrapped around the model that truncates old turns, summarizes at a threshold, offloads to a file, or compacts on a timer. Agent harnesses argued that this loop matters as much as the weights; "the scaffolding gets eaten" argued that most of it is temporary, a crutch a stronger base model will make redundant. A new paper from the University of Washington and Meta FAIR takes the second argument and runs it all the way to the context window.

Their claim, stated as plainly as they state it: give the model unrestricted control over its own context and it beats every human-designed context-management strategy. They call the result a Context Language Model (CLM), and the paper (arXiv 2609.37725, Shao et al., 2026) is an explicit appeal to The Bitter Lesson — stop encoding human priors about how to manage context, and let the model search for a policy.

Overview figure: at top, the context file c_t passes through f_theta^CLM, 'any function of the context', to produce c_t+1, both drawn as files. Four panels below show qualitative behaviors (loops that prune results, reused functions, a subagent tracker), zero-shot results on BrowseComp-Plus and Software World, in-context steering and evolution, and before/after RL bars.
CLMs mirror context as a file the model rewrites arbitrarily, work out of the box with existing LMs, and can be improved in-context or in-weights (Context Language Models paper, Figure 1).

I read the paper, re-derived its serving-cost formula from its own published constants, and cloned the released code to check the implementation. This piece is the mechanism from first principles, then the numbers, labelled by how I know them — reported (the paper's figure, which I did not re-run), measured (computed from the code or a formula), or reasoned (my arithmetic on the other two).

The idea: context as a file

Write a standard language model's turn as a transition on the context ctc_t:

ct+1=ct⊕fθLM(ct)c_{t+1} = c_t \oplus f^{\mathrm{LM}}_{\theta}(c_t)

where ⊕\oplus is concatenation. The model can only ever append: it maps the current context to new tokens, and those tokens are stuck onto the end. Everything else — deciding what in ctc_t is still worth keeping — happens outside θ\theta, in the harness.

A CLM generalizes that one operation:

ct+1=fθCLM(ct)c_{t+1} = f^{\mathrm{CLM}}_{\theta}(c_t)

The next context is now an arbitrary function of the current one, computed by the model. It subsumes append (a CLM can always choose to append), but it also lets the model delete, rewrite in place, reorder, or restructure. The harness stops being the thing that manages context and becomes the thing that gives the model a way to.

The implementation is deliberately unclever, which is the point. The live context is mirrored into a plain text file on disk, and the file's path is handed to the model in its system prompt. The model edits that file with ordinary Bash — sed, a Python one-liner, a re.sub, whatever — exactly as it would edit any other file. The one twist that makes it a context file rather than just a file: after each turn, edits to the mirror are synchronized back into the model's live context and sent to the server for the next turn. When the model does not touch the file, its generated tokens are appended as usual. I checked this against clm/clm_harness/context_env/env.py in the repo: step() writes the mirror, runs the model's command, reconciles the diff, and re-renders the editable context. Edits smaller than 32 tokens are treated as mirror noise and ignored, and the first couple of turns (the system prompt and task) are protected from edits. (measured, from the code.)

Two consequences fall out for free. First, it works zero-shot with existing models — nothing is retrained to make a GPT-5-class or Qwen model into a CLM; you give it a file and a prompt. Second, it extends to multi-agent systems without new machinery: an agent swarm is just several context files coexisting, and a subagent is a file you create and later delete. This is the same move Prime Agent makes with its Python REPL — treat the things a harness usually owns (files, subagents, context) as objects a program manipulates — pushed onto the context itself.

Here is what that looks like across one task. Step through it:

one deep-research task · context as a file vs a fixed harnessillustrative
op 5/5 · answer
Report the recipient and the docid that grounds it.
fixed harness (summarize)docid 58939: lost
answer from a lossy summary
[summary] a 2011 recipient exists; exact id not retained
[answer] name uncertain — docid lost in compaction
live context4.1k tok
peak live 12.1k
cum. compute 3.3k u
CLM (edits the file)docid 58939: kept
answer from the note
[notes] James Gallagher, docid 58939
[answer] James Gallagher (docid 58939)
live context600 tok
peak live 600
cum. compute 1.3k u
step through the taskop 5

Step through the operations. The fixed harness can only append, then compact wholesale when it crosses its 8K budget — so its live context sawtooths up to 12.1k tokens, and the one summary it is forced to take drops the exact docid 58939 the answer needs. The CLM treats the context as a file: it prunes eight search results to one line each, offloads the 9,000-token page to doc1.md behind a pointer, and writes the id into a note on the turn it first sees it. Its live context stays under 600 tokens and the needle never leaves. Cumulative compute (illustrative units) tracks the same gap: re-prefilling a large summarized context costs far more than re-prefilling a short edited one.

The fixed harness has two gears: append, and — when it crosses its budget — compact everything wholesale into a summary. The summary is lossy by construction, so the exact document id the answer needs gets rounded off into a gist. The CLM has a third gear the harness lacks: surgical, in-place edits. It prunes eight verbose search results to one line each with a loop, offloads a 9,000-token page to a file and leaves a one-line pointer, and writes the id into a note on the turn it first sees it. Its live context stays small and keeps the needle. The prose carries the same point the widget does: the win is not "compress harder," it is "keep the right 600 tokens instead of summarizing 12,000 badly."

Why unrestricted beats fixed

To show that fixed strategies fail — not just argue it — the authors built ContextBench, a diagnostic that strips away reasoning and knowledge so that only context management is tested. Four synthetic tasks, each probing one skill:

TaskWhat it testsInput streamMetric
Needle Retentionselective verbatim retention~4K-token chunks, 2–8 needle lines among 140 filler linesneedle lines kept verbatim in the final context
Sudoku Sketchpadin-place surgical editingone move per turn on a 16×16 boardboard versions reproduced exactly
KV Storeoffloading & retrievalbatches of 100 SET ops with random 24-word valuesexact-value accuracy on 24 GET queries
Log Triageoffloading & retrievalbatches of 14–54 service log linesexact-answer accuracy on 24 queries

The twist that makes it honest: every operation arrives as a user message, so it enters the context before the agent can act on it — no harness can truncate it on the way in. The agent controls only when the next operation arrives. Pressure is the ratio of total input to the 32K context limit, pushed up to 24 times the window. The finding is blunt: with GPT-5.4 at a 32K limit, no existing strategy is perfect even here. Summary compaction loses or hallucinates needles; methods without in-place editing have to regenerate the whole Sudoku board for each one-cell move; standard coding tools can offload a value but cannot evict it from the live context on demand. (reported.)

The CLM closes those gaps because it is not restricted to a menu of actions. And left to edit freely, it does things no one wrote into a harness. The paper catalogs them from the zero-shot runs: on a multi-agent orchestration task it maintains an in-context scoreboard through 163 in-place edits while holding the context at 6–8K tokens; it invents a new notes chat role for its own internal state; it writes loops to strip irrelevant search results; and in one run it defines a compact_turns helper and calls it 37 times to compress observations behind a progress note. (reported.) These are the "emergent strategies" the bitter-lesson framing predicts: search finds policies a human designer would not have specified.

The receipts, zero-shot

The headline is that a CLM, applied to existing models with no training, beats specialized harnesses on long-horizon tasks — and does it at lower cost. Cost here is prefix-reuse FLOPs, the paper's accounting of real serving compute, which I come back to below. On deep research and coding, with Qwen3.6-27B at a 32K budget:

Three scatter panels, accuracy versus prefix-reuse PFLOPs per question. On BrowseComp-Plus, TerminalBench 2.1 and TBLite, the CLM point sits up and to the left of the baselines Summary, MEM1, Self-Compact, ACM, RLM and Base, on the Pareto frontier.
CLMs land on the accuracy-vs-cost Pareto frontier against action-based and harness-defined baselines, with Qwen3.6-27B at a 32K context limit (Context Language Models paper, Figure 5).

On open-ended discovery, the horizons stretch from hours to a full day. On four math-optimization problems from the AlphaEvolve/OpenEvolve suite, a general CLM agent with a minimal Bash interface — just handed the evolutionary algorithm as in-context guidance — beats the specialized OpenEvolve workflow on all four, using Claude 4.6 Sonnet:

ProblemOpenEvolveCLMdirection
Circle packing2.5412.618higher is better
Heilbronn triangle0.031270.03653higher is better
Min-max / min-dist0.076900.07758higher is better
Erdős min-overlap0.381230.38094lower is better

The circle-packing margin is 2.618/2.541=1.0302.618 / 2.541 = 1.030, a 3.0% gain; Heilbronn is 0.03653/0.03127=1.1680.03653 / 0.03127 = 1.168, 16.8%. (reasoned, matching the paper's "up to 16.8% and 3.0%" — reported.) That a general agent with "less fixed orchestration" tops a workflow built specifically for evolutionary program search is the cleanest statement of the thesis in the paper.

The two longest runs are the most striking. On EdgeBench — optimize a single repository for up to 12 hours — CLM with Qwen3.6-27B reaches 44.6 at 179 prefix-reuse PFLOPs per trial, versus 42.3 at 437 PFLOPs for Codex-style summarization. That is the abstract's "5% higher scores with 59% fewer FLOPs": 44.6/42.3=1.05444.6/42.3 = 1.054 and (437−179)/437=0.590(437-179)/437 = 0.590. (reasoned, from reported numbers.) On Software World — six agents jointly optimizing six interdependent repositories over 24 hours, scored on four unseen downstream packages — the CLM swarm delivers 65% greater downstream speedup than a summary-based swarm at the same spend. (reported.) This is the result that most resists a "benchmark artifact" read, because the score is measured on packages the agents never touched — the agent-swarm argument with coordination moved inside the model.

Learning the policy: in context, then in weights

Because context management is now a model behavior, it can be taught like one — first with words, then with gradients.

Steering is the cheap version. A single sentence appended to the prompt — "compact once you reach Y tokens," "compact at sub-question boundaries," "back up the context before compacting" — changes the policy, with no harness change and no new weights. The steering panel in Figure 1 shows the instructed compaction threshold tracking the requested one almost on the x=yx = y line.

In-context evolution is steering with a search loop around it: the agent runs on a training split, a proposer model reads the traces and writes a better skill document, you keep the winner, and evaluate the final skill once on a held-out split. On ContextBench's KV Store task at 32K, starting from no instruction, the evolved skill raises held-out accuracy from 38.3% to 74.2% — a 35.9-point gain (74.2−38.3=35.974.2 - 38.3 = 35.9) — at lower compute. (reasoned, matching the reported "35.9 points.")

Reinforcement learning internalizes the strategy into the weights. Because context edits change the input across turns, they use stepwise GRPO: sample a group of full trajectories, compute the usual outcome advantage, and assign it to every model call in the trajectory. The new piece is a success-gated efficiency advantage. The problem it solves: rewarding "fewer tokens" or "more edits" directly invites reward hacking — the model learns to make pointless edits that throw away information or wreck prefix reuse. So the efficiency signal only ranks among already-successful trajectories. For successful trajectory ii in group gg with prefix-reuse FLOPs cic_i and group mean cˉg\bar{c}_g:

Aieff=clip⁡ ⁣(cˉg−cicˉg, −1, 1)A_i^{\mathrm{eff}} = \operatorname{clip}\!\left(\frac{\bar{c}_g - c_i}{\bar{c}_g},\, -1,\, 1\right)

and the total advantage is Ai=Aiout+weffAieffA_i = A_i^{\mathrm{out}} + w_{\mathrm{eff}} A_i^{\mathrm{eff}}. Failed trajectories, and groups with fewer than two successes, get Aieff=0A_i^{\mathrm{eff}} = 0. Efficiency never trades away correctness; it only breaks ties between correct runs by cost.

Training Qwen3.5-9B on OpenResearcher and evaluating on held-out BrowseComp-Plus:

Acc. (%)PFLOPs / Q
Summary harness34.7 → 42.14.01 → 2.19
CLM28.8 → 42.51.52 → 1.34

Before RL, the small CLM sits at 28.8%, six points behind the summary harness — a 9B model is not a great context manager out of the box. After RL it gains 13.7 points to 42.5%, the abstract's 47.6% relative improvement (42.5/28.8=1.47642.5/28.8 = 1.476), while its per-question cost falls from 1.52 to 1.34 PFLOPs — 12% fewer ((1.52−1.34)/1.52=0.118(1.52-1.34)/1.52 = 0.118). That jump from 28.8 to 42.5 is the strategy being learned in the weights. (reasoned, from the reported table.) It ends 0.4 points above the trained summary harness while spending 38.8% fewer FLOPs than it ((2.19−1.34)/2.19=0.388(2.19-1.34)/2.19 = 0.388). (reasoned.) The policy that a human would have hard-coded into a harness is instead discovered and folded into the weights.

What arbitrary edits cost to serve

There is a tax, and the paper is honest about it. Modern servers make long contexts affordable through prefix caching: the KV cache is content-addressed, so a turn that only appends reuses everything computed before and prefills just the new tokens. (If you want the why, self-attention ties each token's key and value to every token before it and, through rotary position encodings, to its absolute position.) An edit in the middle breaks that: every token after the first mismatch has to be re-prefilled, including text that did not change.

To make the cost legible, the authors define prefix-reuse FLOPs and work it out for Qwen3.6-27B, a hybrid model with 48 of its 64 layers using linear attention. Per-token projection cost is Ctoken=48.70×109C_{\mathrm{token}} = 48.70 \times 10^{9} FLOPs and each query-key pair costs Cattn=3.93×105C_{\mathrm{attn}} = 3.93 \times 10^{5}. For a turn with prompt length PP, generated length GG, and a reusable prefix of RR tokens, the cost is

F=Ctoken(P−R+G)+Cattn ⁣[12(P2−R2)+GP+12G2]F = C_{\mathrm{token}}(P - R + G) + C_{\mathrm{attn}}\!\left[\tfrac{1}{2}(P^2 - R^2) + GP + \tfrac{1}{2}G^2\right]

I re-implemented this formula from the constants above and confirmed it reproduces the paper's three worked points exactly. Drag the edit position:

one Qwen3.6-27B turn · P = 20,000, G = 500 · prefix-reuse FLOPspaper formula
reused prefix: 50%re-prefilled: 50%
cache reused
re-prefill
5.74x10^14
FLOPs this turn
4.1x
of an append-only turn
reusable prefix R (drag the edit point)R = 10,000 tok

Appending to the context reuses 18,000 of the 20,000 prompt tokens and costs 1.41 x10^14 FLOPs — prefix caching avoids 87% of what the turn would need cold. An edit in the middle leaves 10,000 reusable tokens and the cost climbs to 5.74 x10^14. An edit at the start reuses nothing and costs 10.81 x10^14 — 7.7x the append-only turn. That re-prefill tax is the price of arbitrary edits, and it is exactly what prefix-reuse FLOPs measure and Suffix Cache Reuse attacks.

For an illustrative turn with P=20,000P = 20{,}000 and G=500G = 500: appending (reusing 18,000 tokens) costs 1.41×10141.41 \times 10^{14} FLOPs, and prefix caching has avoided 87% of what a cold turn would need. An edit in the middle that leaves 10,000 reusable tokens raises it to 5.74×10145.74 \times 10^{14}. An edit at the very start reuses nothing and costs 10.81×101410.81 \times 10^{14} — 7.7 times the append-only turn (10.81/1.41=7.6710.81/1.41 = 7.67). (measured; the formula reproduces the paper's reported 1.411.41, 5.745.74, 10.8110.81 and 7.7×.) The important thing is that the paper reports all its efficiency numbers under this metric, under standard serving — so the "fewer FLOPs" wins above already pay the edit tax. CLMs come out ahead not by avoiding re-prefills but by keeping the context so small that each re-prefill is cheap.

Diagram. Original context is three blocks A, B, C. An edit replaces B with B-prime. Standard serving reuses the cache for A, prefills B-prime, and re-prefills all of C. Suffix Cache Reuse reuses A, prefills B-prime, and reuses the cache for C as well.
Suffix Cache Reuse relocates the surviving suffix's cached states past the edit instead of re-prefilling them (Context Language Models paper, Figure 4).

Then they attack the tax directly with Suffix Cache Reuse (SCR), a patch to SGLang. When an edit replaces BB with B′B' in a context [A B C][A\,B\,C], standard serving re-prefills B′B' and the surviving suffix CC. SCR instead reuses CC's cached keys and values: it diffs the new prompt against the old, finds the surviving spans, re-rotates their rotary position encodings to the new positions, and splices them in after B′B'. Only B′B' is prefilled. It relocates up to KK spans per edit, largest first; the released config defaults to K=6 (KVREUSE_MAX_BLOCKS=6), which I confirmed in suffix_cache_reuse/config.py. (measured, from the code.) The reused states are slightly stale — computed under the pre-edit context — so SCR approximates re-prefilling, and the cap bounds how much approximation any one edit introduces.

The result: SCR matches standard SGLang at 65.0% of its prefix-reuse FLOPs on BrowseComp-Plus — a 35% reduction in server-side compute at matched performance. (reported.) A nice side effect: reasoning models often strip earlier turns' reasoning blocks before the next turn, which also forces a re-prefill, and SCR fixes that too even when the model never edits its own context. Of the 7.8% of prompt tokens SCR reuses beyond prefix-cache hits, 5.3 points come from reasoning-stripping and only 2.5 from edits — so most of the benefit applies to ordinary reasoning-model serving. (reported.)

Where it is still a promise

The thesis is strong and the receipts are mostly real, but a few things deserve flags.

None of that undercuts the core finding, which I think is right: context management was the last big piece of the agent loop still living entirely in hand-written scaffolding, and this is a credible argument — with numbers under an honest cost metric — that it does not have to. The harness gets eaten; the context goes first.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Context Language Models: the bitter lesson comes for context management", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026contextlanguagemodels,
  author = {Satyajit Ghana},
  title  = {Context Language Models: the bitter lesson comes for context management},
  url    = {https://ai.thesatyajit.com/articles/context-language-models},
  year   = {2026}
}
share