2026-10-02 · 16 min · explainer · agents · llm · long-context · inference-optimization · kv-cache · reinforcement-learning
A 1:42 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.
› transcript
Hi, I'm Vesper! Let me show you how a model can run its own memory, by treating its context as a file it rewrites. Instead of a harness deciding what to forget, the model edits its context like a file: append, rewrite a line in place, or offload a chunk. A turn starts with the live context. It is mirrored as a plain text file the model can open. The model reads the file and edits it with Bash. It can append, rewrite a line in place, or move a chunk out to storage. The edit syncs back and becomes the context for the next turn. A fixed harness summarizes on a schedule. One lossy gist, and the exact fact you needed is gone. The model instead keeps the right lines and drops the rest, in any strategy it can write as code. With no training, this beats the strongest harness on a deep research benchmark, and gets there using less compute. Online reinforcement learning folds the strategy into the weights. A small model jumps a lot, and spends less compute doing it. An arbitrary edit is not free to serve. Standard serving re-prefills the whole tail after the edit, even text that never changed. Suffix Cache Reuse keeps the unchanged tail's cache and splices it back, so only the edited part is recomputed. Context management was the last piece of the agent loop written by hand. The model can learn to do it better. Context is a file. The model reads it, edits it, and keeps what matters. That's the whole idea. Thanks for watching! Every source is in the full article. I'm Vesper. Bye!
Every long-running agent eventually runs out of room. The transcript grows, the context window fills, and something has to decide what to throw away. Today that something is almost never the model. It is the harness — the loop wrapped around the model that truncates old turns, summarizes at a threshold, offloads to a file, or compacts on a timer. Agent harnesses argued that this loop matters as much as the weights; "the scaffolding gets eaten" argued that most of it is temporary, a crutch a stronger base model will make redundant. A new paper from the University of Washington and Meta FAIR takes the second argument and runs it all the way to the context window.
Their claim, stated as plainly as they state it: give the model unrestricted control over its own context and it beats every human-designed context-management strategy. They call the result a Context Language Model (CLM), and the paper (arXiv 2609.37725, Shao et al., 2026) is an explicit appeal to The Bitter Lesson — stop encoding human priors about how to manage context, and let the model search for a policy.

I read the paper, re-derived its serving-cost formula from its own published constants, and cloned the released code to check the implementation. This piece is the mechanism from first principles, then the numbers, labelled by how I know them — reported (the paper's figure, which I did not re-run), measured (computed from the code or a formula), or reasoned (my arithmetic on the other two).
The idea: context as a file
Write a standard language model's turn as a transition on the context :
where is concatenation. The model can only ever append: it maps the current context to new tokens, and those tokens are stuck onto the end. Everything else — deciding what in is still worth keeping — happens outside , in the harness.
A CLM generalizes that one operation:
The next context is now an arbitrary function of the current one, computed by the model. It subsumes append (a CLM can always choose to append), but it also lets the model delete, rewrite in place, reorder, or restructure. The harness stops being the thing that manages context and becomes the thing that gives the model a way to.
The implementation is deliberately unclever, which is the point. The live context is mirrored into a plain text file on disk, and the file's path is handed to the model in its system prompt. The model edits that file with ordinary Bash — sed, a Python one-liner, a re.sub, whatever — exactly as it would edit any other file. The one twist that makes it a context file rather than just a file: after each turn, edits to the mirror are synchronized back into the model's live context and sent to the server for the next turn. When the model does not touch the file, its generated tokens are appended as usual. I checked this against clm/clm_harness/context_env/env.py in the repo: step() writes the mirror, runs the model's command, reconciles the diff, and re-renders the editable context. Edits smaller than 32 tokens are treated as mirror noise and ignored, and the first couple of turns (the system prompt and task) are protected from edits. (measured, from the code.)
Two consequences fall out for free. First, it works zero-shot with existing models — nothing is retrained to make a GPT-5-class or Qwen model into a CLM; you give it a file and a prompt. Second, it extends to multi-agent systems without new machinery: an agent swarm is just several context files coexisting, and a subagent is a file you create and later delete. This is the same move Prime Agent makes with its Python REPL — treat the things a harness usually owns (files, subagents, context) as objects a program manipulates — pushed onto the context itself.
Here is what that looks like across one task. Step through it:
Step through the operations. The fixed harness can only append, then compact wholesale when it crosses its 8K budget — so its live context sawtooths up to 12.1k tokens, and the one summary it is forced to take drops the exact docid 58939 the answer needs. The CLM treats the context as a file: it prunes eight search results to one line each, offloads the 9,000-token page to doc1.md behind a pointer, and writes the id into a note on the turn it first sees it. Its live context stays under 600 tokens and the needle never leaves. Cumulative compute (illustrative units) tracks the same gap: re-prefilling a large summarized context costs far more than re-prefilling a short edited one.
The fixed harness has two gears: append, and — when it crosses its budget — compact everything wholesale into a summary. The summary is lossy by construction, so the exact document id the answer needs gets rounded off into a gist. The CLM has a third gear the harness lacks: surgical, in-place edits. It prunes eight verbose search results to one line each with a loop, offloads a 9,000-token page to a file and leaves a one-line pointer, and writes the id into a note on the turn it first sees it. Its live context stays small and keeps the needle. The prose carries the same point the widget does: the win is not "compress harder," it is "keep the right 600 tokens instead of summarizing 12,000 badly."
Why unrestricted beats fixed
To show that fixed strategies fail — not just argue it — the authors built ContextBench, a diagnostic that strips away reasoning and knowledge so that only context management is tested. Four synthetic tasks, each probing one skill:
| Task | What it tests | Input stream | Metric |
|---|---|---|---|
| Needle Retention | selective verbatim retention | ~4K-token chunks, 2–8 needle lines among 140 filler lines | needle lines kept verbatim in the final context |
| Sudoku Sketchpad | in-place surgical editing | one move per turn on a 16×16 board | board versions reproduced exactly |
| KV Store | offloading & retrieval | batches of 100 SET ops with random 24-word values | exact-value accuracy on 24 GET queries |
| Log Triage | offloading & retrieval | batches of 14–54 service log lines | exact-answer accuracy on 24 queries |
The twist that makes it honest: every operation arrives as a user message, so it enters the context before the agent can act on it — no harness can truncate it on the way in. The agent controls only when the next operation arrives. Pressure is the ratio of total input to the 32K context limit, pushed up to 24 times the window. The finding is blunt: with GPT-5.4 at a 32K limit, no existing strategy is perfect even here. Summary compaction loses or hallucinates needles; methods without in-place editing have to regenerate the whole Sudoku board for each one-cell move; standard coding tools can offload a value but cannot evict it from the live context on demand. (reported.)
The CLM closes those gaps because it is not restricted to a menu of actions. And left to edit freely, it does things no one wrote into a harness. The paper catalogs them from the zero-shot runs: on a multi-agent orchestration task it maintains an in-context scoreboard through 163 in-place edits while holding the context at 6–8K tokens; it invents a new notes chat role for its own internal state; it writes loops to strip irrelevant search results; and in one run it defines a compact_turns helper and calls it 37 times to compress observations behind a progress note. (reported.) These are the "emergent strategies" the bitter-lesson framing predicts: search finds policies a human designer would not have specified.
The receipts, zero-shot
The headline is that a CLM, applied to existing models with no training, beats specialized harnesses on long-horizon tasks — and does it at lower cost. Cost here is prefix-reuse FLOPs, the paper's accounting of real serving compute, which I come back to below. On deep research and coding, with Qwen3.6-27B at a 32K budget:

- BrowseComp-Plus (deep research): CLM scores 59.4%, beating the strongest baseline — Codex-style summarization — by 11.4% relative, while using 21.5% fewer prefix-reuse FLOPs than it (and 28.9% fewer than the next method, MEM1). (reported.)
- TerminalBench 2.1 (terminal coding): CLM matches Codex-style summarization's accuracy using only 70% of its FLOPs. (reported.)
- TBLite: CLM beats it, 73.7% against 67.0%, at 91% of its FLOPs. (reported.)
On open-ended discovery, the horizons stretch from hours to a full day. On four math-optimization problems from the AlphaEvolve/OpenEvolve suite, a general CLM agent with a minimal Bash interface — just handed the evolutionary algorithm as in-context guidance — beats the specialized OpenEvolve workflow on all four, using Claude 4.6 Sonnet:
| Problem | OpenEvolve | CLM | direction |
|---|---|---|---|
| Circle packing | 2.541 | 2.618 | higher is better |
| Heilbronn triangle | 0.03127 | 0.03653 | higher is better |
| Min-max / min-dist | 0.07690 | 0.07758 | higher is better |
| Erdős min-overlap | 0.38123 | 0.38094 | lower is better |
The circle-packing margin is , a 3.0% gain; Heilbronn is , 16.8%. (reasoned, matching the paper's "up to 16.8% and 3.0%" — reported.) That a general agent with "less fixed orchestration" tops a workflow built specifically for evolutionary program search is the cleanest statement of the thesis in the paper.
The two longest runs are the most striking. On EdgeBench — optimize a single repository for up to 12 hours — CLM with Qwen3.6-27B reaches 44.6 at 179 prefix-reuse PFLOPs per trial, versus 42.3 at 437 PFLOPs for Codex-style summarization. That is the abstract's "5% higher scores with 59% fewer FLOPs": and . (reasoned, from reported numbers.) On Software World — six agents jointly optimizing six interdependent repositories over 24 hours, scored on four unseen downstream packages — the CLM swarm delivers 65% greater downstream speedup than a summary-based swarm at the same spend. (reported.) This is the result that most resists a "benchmark artifact" read, because the score is measured on packages the agents never touched — the agent-swarm argument with coordination moved inside the model.
Learning the policy: in context, then in weights
Because context management is now a model behavior, it can be taught like one — first with words, then with gradients.
Steering is the cheap version. A single sentence appended to the prompt — "compact once you reach Y tokens," "compact at sub-question boundaries," "back up the context before compacting" — changes the policy, with no harness change and no new weights. The steering panel in Figure 1 shows the instructed compaction threshold tracking the requested one almost on the line.
In-context evolution is steering with a search loop around it: the agent runs on a training split, a proposer model reads the traces and writes a better skill document, you keep the winner, and evaluate the final skill once on a held-out split. On ContextBench's KV Store task at 32K, starting from no instruction, the evolved skill raises held-out accuracy from 38.3% to 74.2% — a 35.9-point gain () — at lower compute. (reasoned, matching the reported "35.9 points.")
Reinforcement learning internalizes the strategy into the weights. Because context edits change the input across turns, they use stepwise GRPO: sample a group of full trajectories, compute the usual outcome advantage, and assign it to every model call in the trajectory. The new piece is a success-gated efficiency advantage. The problem it solves: rewarding "fewer tokens" or "more edits" directly invites reward hacking — the model learns to make pointless edits that throw away information or wreck prefix reuse. So the efficiency signal only ranks among already-successful trajectories. For successful trajectory in group with prefix-reuse FLOPs and group mean :
and the total advantage is . Failed trajectories, and groups with fewer than two successes, get . Efficiency never trades away correctness; it only breaks ties between correct runs by cost.
Training Qwen3.5-9B on OpenResearcher and evaluating on held-out BrowseComp-Plus:
| Acc. (%) | PFLOPs / Q | |
|---|---|---|
| Summary harness | 34.7 → 42.1 | 4.01 → 2.19 |
| CLM | 28.8 → 42.5 | 1.52 → 1.34 |
Before RL, the small CLM sits at 28.8%, six points behind the summary harness — a 9B model is not a great context manager out of the box. After RL it gains 13.7 points to 42.5%, the abstract's 47.6% relative improvement (), while its per-question cost falls from 1.52 to 1.34 PFLOPs — 12% fewer (). That jump from 28.8 to 42.5 is the strategy being learned in the weights. (reasoned, from the reported table.) It ends 0.4 points above the trained summary harness while spending 38.8% fewer FLOPs than it (). (reasoned.) The policy that a human would have hard-coded into a harness is instead discovered and folded into the weights.
What arbitrary edits cost to serve
There is a tax, and the paper is honest about it. Modern servers make long contexts affordable through prefix caching: the KV cache is content-addressed, so a turn that only appends reuses everything computed before and prefills just the new tokens. (If you want the why, self-attention ties each token's key and value to every token before it and, through rotary position encodings, to its absolute position.) An edit in the middle breaks that: every token after the first mismatch has to be re-prefilled, including text that did not change.
To make the cost legible, the authors define prefix-reuse FLOPs and work it out for Qwen3.6-27B, a hybrid model with 48 of its 64 layers using linear attention. Per-token projection cost is FLOPs and each query-key pair costs . For a turn with prompt length , generated length , and a reusable prefix of tokens, the cost is
I re-implemented this formula from the constants above and confirmed it reproduces the paper's three worked points exactly. Drag the edit position:
Appending to the context reuses 18,000 of the 20,000 prompt tokens and costs 1.41 x10^14 FLOPs — prefix caching avoids 87% of what the turn would need cold. An edit in the middle leaves 10,000 reusable tokens and the cost climbs to 5.74 x10^14. An edit at the start reuses nothing and costs 10.81 x10^14 — 7.7x the append-only turn. That re-prefill tax is the price of arbitrary edits, and it is exactly what prefix-reuse FLOPs measure and Suffix Cache Reuse attacks.
For an illustrative turn with and : appending (reusing 18,000 tokens) costs FLOPs, and prefix caching has avoided 87% of what a cold turn would need. An edit in the middle that leaves 10,000 reusable tokens raises it to . An edit at the very start reuses nothing and costs — 7.7 times the append-only turn (). (measured; the formula reproduces the paper's reported , , and 7.7×.) The important thing is that the paper reports all its efficiency numbers under this metric, under standard serving — so the "fewer FLOPs" wins above already pay the edit tax. CLMs come out ahead not by avoiding re-prefills but by keeping the context so small that each re-prefill is cheap.

Then they attack the tax directly with Suffix Cache Reuse (SCR), a patch to SGLang. When an edit replaces with in a context , standard serving re-prefills and the surviving suffix . SCR instead reuses 's cached keys and values: it diffs the new prompt against the old, finds the surviving spans, re-rotates their rotary position encodings to the new positions, and splices them in after . Only is prefilled. It relocates up to spans per edit, largest first; the released config defaults to K=6 (KVREUSE_MAX_BLOCKS=6), which I confirmed in suffix_cache_reuse/config.py. (measured, from the code.) The reused states are slightly stale — computed under the pre-edit context — so SCR approximates re-prefilling, and the cap bounds how much approximation any one edit introduces.
The result: SCR matches standard SGLang at 65.0% of its prefix-reuse FLOPs on BrowseComp-Plus — a 35% reduction in server-side compute at matched performance. (reported.) A nice side effect: reasoning models often strip earlier turns' reasoning blocks before the next turn, which also forces a re-prefill, and SCR fixes that too even when the model never edits its own context. Of the 7.8% of prompt tokens SCR reuses beyond prefix-cache hits, 5.3 points come from reasoning-stripping and only 2.5 from edits — so most of the benefit applies to ordinary reasoning-model serving. (reported.)
Where it is still a promise
The thesis is strong and the receipts are mostly real, but a few things deserve flags.
- I verified arithmetic and code structure, not the runs. Every benchmark number here is reported — the training and 24-hour evaluations are not reproducible in a sandbox, and I did not re-run them. What I independently checked is the serving-cost formula (it reproduces the paper's figures from its own constants), the context-as-a-file mechanism (
ContextEnvreally does mirror, edit, and sync), and the SCR defaults. Treat the benchmark wins as the authors' measurements, not mine. - ContextBench is not released yet. The repo's README lists it as "coming soon," so the one benchmark built specifically to isolate context management — and the source of the 35.9-point gain — cannot be independently run today. The harness, the RL code, and SCR are released.
- The license is research-only. The code is CC BY-NC 4.0, not a permissive license. Fine for study; not something you drop into a commercial agent.
- Context-length awareness is a real limitation. To decide when to edit, a model has to know how full its context is, and the paper's own appendix shows existing LMs are poor at this — they guess bucketed values rather than reading their true length, and need external hints to calibrate. A CLM that cannot feel its own budget leans back on the harness to tell it — exactly the dependency the paper is trying to remove.
None of that undercuts the core finding, which I think is right: context management was the last big piece of the agent loop still living entirely in hand-written scaffolding, and this is a credible argument — with numbers under an honest cost metric — that it does not have to. The harness gets eaten; the context goes first.