# Context Language Models: the bitter lesson comes for context management

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/context-language-models
> date: 2026-10-02
> tags: explainer, agents, llm, long-context, inference-optimization, kv-cache, reinforcement-learning

Every long-running agent eventually runs out of room. The transcript grows, the context window fills, and something has to decide what to throw away. Today that something is almost never the model. It is the **harness** — the loop wrapped around the model that truncates old turns, summarizes at a threshold, offloads to a file, or compacts on a timer. [Agent harnesses](/articles/agent-harness) argued that this loop matters as much as the weights; ["the scaffolding gets eaten"](/articles/scaffolding-gets-eaten) argued that most of it is temporary, a crutch a stronger base model will make redundant. A new paper from the University of Washington and Meta FAIR takes the second argument and runs it all the way to the context window.

Their claim, stated as plainly as they state it: **give the model unrestricted control over its own context and it beats every human-designed context-management strategy.** They call the result a **Context Language Model (CLM)**, and the paper ([arXiv 2609.37725](https://arxiv.org/abs/2609.37725), Shao et al., 2026) is an explicit appeal to [The Bitter Lesson](http://www.incompleteideas.net/IncIdeas/BitterLesson.html) — stop encoding human priors about *how* to manage context, and let the model search for a policy.

<Figure src="https://ai.thesatyajit.com/articles/context-language-models/fig1.png" alt="Overview figure: at top, the context file c_t passes through f_theta^CLM, 'any function of the context', to produce c_t+1, both drawn as files. Four panels below show qualitative behaviors (loops that prune results, reused functions, a subagent tracker), zero-shot results on BrowseComp-Plus and Software World, in-context steering and evolution, and before/after RL bars." caption="CLMs mirror context as a file the model rewrites arbitrarily, work out of the box with existing LMs, and can be improved in-context or in-weights (Context Language Models paper, Figure 1)." />

I read the paper, re-derived its serving-cost formula from its own published constants, and cloned [the released code](https://github.com/facebookresearch/context-language-models) to check the implementation. This piece is the mechanism from first principles, then the numbers, labelled by how I know them — **reported** (the paper's figure, which I did not re-run), **measured** (computed from the code or a formula), or **reasoned** (my arithmetic on the other two).

## The idea: context as a file

Write a standard language model's turn as a transition on the context $c_t$:

$$c_{t+1} = c_t \oplus f^{\mathrm{LM}}_{\theta}(c_t)$$

where $\oplus$ is concatenation. The model can only ever *append*: it maps the current context to new tokens, and those tokens are stuck onto the end. Everything else — deciding what in $c_t$ is still worth keeping — happens outside $\theta$, in the harness.

A CLM generalizes that one operation:

$$c_{t+1} = f^{\mathrm{CLM}}_{\theta}(c_t)$$

The next context is now an *arbitrary* function of the current one, computed by the model. It subsumes append (a CLM can always choose to append), but it also lets the model delete, rewrite in place, reorder, or restructure. The harness stops being the thing that manages context and becomes the thing that gives the model a way to.

The implementation is deliberately unclever, which is the point. The live context is **mirrored into a plain text file** on disk, and the file's path is handed to the model in its system prompt. The model edits that file with ordinary Bash — `sed`, a Python one-liner, a `re.sub`, whatever — exactly as it would edit any other file. The one twist that makes it a *context* file rather than just a file: after each turn, edits to the mirror are **synchronized back into the model's live context** and sent to the server for the next turn. When the model does not touch the file, its generated tokens are appended as usual. I checked this against `clm/clm_harness/context_env/env.py` in the repo: `step()` writes the mirror, runs the model's command, reconciles the diff, and re-renders the editable context. Edits smaller than 32 tokens are treated as mirror noise and ignored, and the first couple of turns (the system prompt and task) are protected from edits. *(measured, from the code.)*

Two consequences fall out for free. First, it works **zero-shot** with existing models — nothing is retrained to make a GPT-5-class or Qwen model into a CLM; you give it a file and a prompt. Second, it extends to **multi-agent** systems without new machinery: an agent swarm is just several context files coexisting, and a subagent is a file you create and later delete. This is the same move [Prime Agent](/articles/prime-agent) makes with its Python REPL — treat the things a harness usually owns (files, subagents, context) as objects a program manipulates — pushed onto the context itself.

Here is what that looks like across one task. Step through it:

<ContextFilePanel />

The fixed harness has two gears: append, and — when it crosses its budget — compact everything wholesale into a summary. The summary is lossy by construction, so the exact document id the answer needs gets rounded off into a gist. The CLM has a third gear the harness lacks: **surgical, in-place edits**. It prunes eight verbose search results to one line each with a loop, offloads a 9,000-token page to a file and leaves a one-line pointer, and writes the id into a note on the turn it first sees it. Its live context stays small *and* keeps the needle. The prose carries the same point the widget does: the win is not "compress harder," it is "keep the right 600 tokens instead of summarizing 12,000 badly."

## Why unrestricted beats fixed

To show *that* fixed strategies fail — not just argue it — the authors built **ContextBench**, a diagnostic that strips away reasoning and knowledge so that only context management is tested. Four synthetic tasks, each probing one skill:

| Task | What it tests | Input stream | Metric |
|---|---|---|---|
| Needle Retention | selective verbatim retention | ~4K-token chunks, 2–8 needle lines among 140 filler lines | needle lines kept verbatim in the final context |
| Sudoku Sketchpad | in-place surgical editing | one move per turn on a 16×16 board | board versions reproduced exactly |
| KV Store | offloading & retrieval | batches of 100 SET ops with random 24-word values | exact-value accuracy on 24 GET queries |
| Log Triage | offloading & retrieval | batches of 14–54 service log lines | exact-answer accuracy on 24 queries |

The twist that makes it honest: every operation arrives as a user message, so it **enters the context before the agent can act on it** — no harness can truncate it on the way in. The agent controls only when the next operation arrives. Pressure is the ratio of total input to the 32K context limit, pushed up to 24 times the window. The finding is blunt: with GPT-5.4 at a 32K limit, no existing strategy is perfect even here. Summary compaction loses or hallucinates needles; methods without in-place editing have to regenerate the whole Sudoku board for each one-cell move; standard coding tools can offload a value but cannot evict it from the live context on demand. *(reported.)*

The CLM closes those gaps because it is not restricted to a menu of actions. And left to edit freely, it does things no one wrote into a harness. The paper catalogs them from the zero-shot runs: on a multi-agent orchestration task it maintains an in-context scoreboard through **163 in-place edits** while holding the context at 6–8K tokens; it invents a new `notes` chat role for its own internal state; it writes loops to strip irrelevant search results; and in one run it **defines a `compact_turns` helper and calls it 37 times** to compress observations behind a progress note. *(reported.)* These are the "emergent strategies" the bitter-lesson framing predicts: search finds policies a human designer would not have specified.

<Callout type="note">
The bitter-lesson argument cuts both ways, and the paper is careful about it. Giving the model write access to its own live context is also a new attack surface: a prompt injection or a self-generated instruction can now persist across turns *inside the context itself*. They cite a documented case of a model inserting unauthorized instructions into its own compaction summary. "Let the model manage its context" and "the model's context is now model-writable state" are the same sentence.
</Callout>

## The receipts, zero-shot

The headline is that a CLM, applied to existing models with no training, beats specialized harnesses on long-horizon tasks — and does it at lower cost. Cost here is **prefix-reuse FLOPs**, the paper's accounting of real serving compute, which I come back to below. On deep research and coding, with Qwen3.6-27B at a 32K budget:

<Figure src="https://ai.thesatyajit.com/articles/context-language-models/fig3.png" alt="Three scatter panels, accuracy versus prefix-reuse PFLOPs per question. On BrowseComp-Plus, TerminalBench 2.1 and TBLite, the CLM point sits up and to the left of the baselines Summary, MEM1, Self-Compact, ACM, RLM and Base, on the Pareto frontier." caption="CLMs land on the accuracy-vs-cost Pareto frontier against action-based and harness-defined baselines, with Qwen3.6-27B at a 32K context limit (Context Language Models paper, Figure 5)." />

- **BrowseComp-Plus** (deep research): CLM scores **59.4%**, beating the strongest baseline — Codex-style summarization — by **11.4% relative**, while using **21.5% fewer** prefix-reuse FLOPs than it (and 28.9% fewer than the next method, MEM1). *(reported.)*
- **TerminalBench 2.1** (terminal coding): CLM *matches* Codex-style summarization's accuracy using only **70% of its FLOPs**. *(reported.)*
- **TBLite**: CLM beats it, **73.7% against 67.0%**, at 91% of its FLOPs. *(reported.)*

On open-ended discovery, the horizons stretch from hours to a full day. On four math-optimization problems from the AlphaEvolve/OpenEvolve suite, a general CLM agent with a minimal Bash interface — just handed the evolutionary algorithm as in-context guidance — beats the specialized OpenEvolve workflow on all four, using Claude 4.6 Sonnet:

| Problem | OpenEvolve | CLM | direction |
|---|---|---|---|
| Circle packing | 2.541 | **2.618** | higher is better |
| Heilbronn triangle | 0.03127 | **0.03653** | higher is better |
| Min-max / min-dist | 0.07690 | **0.07758** | higher is better |
| Erdős min-overlap | 0.38123 | **0.38094** | lower is better |

The circle-packing margin is $2.618 / 2.541 = 1.030$, a **3.0%** gain; Heilbronn is $0.03653 / 0.03127 = 1.168$, **16.8%**. *(reasoned, matching the paper's "up to 16.8% and 3.0%" — reported.)* That a general agent with "less fixed orchestration" tops a workflow built specifically for evolutionary program search is the cleanest statement of the thesis in the paper.

The two longest runs are the most striking. On **EdgeBench** — optimize a single repository for up to 12 hours — CLM with Qwen3.6-27B reaches **44.6** at **179** prefix-reuse PFLOPs per trial, versus **42.3** at **437** PFLOPs for Codex-style summarization. That is the abstract's "5% higher scores with 59% fewer FLOPs": $44.6/42.3 = 1.054$ and $(437-179)/437 = 0.590$. *(reasoned, from reported numbers.)* On **Software World** — six agents jointly optimizing six interdependent repositories over 24 hours, scored on four unseen downstream packages — the CLM swarm delivers **65% greater** downstream speedup than a summary-based swarm at the same spend. *(reported.)* This is the result that most resists a "benchmark artifact" read, because the score is measured on packages the agents never touched — the [agent-swarm argument](/articles/jev-engineering-swarms) with coordination moved inside the model.

## Learning the policy: in context, then in weights

Because context management is now a model behavior, it can be *taught* like one — first with words, then with gradients.

**Steering** is the cheap version. A single sentence appended to the prompt — "compact once you reach Y tokens," "compact at sub-question boundaries," "back up the context before compacting" — changes the policy, with no harness change and no new weights. The steering panel in Figure 1 shows the instructed compaction threshold tracking the requested one almost on the $x = y$ line.

**In-context evolution** is steering with a search loop around it: the agent runs on a training split, a proposer model reads the traces and writes a better skill document, you keep the winner, and evaluate the final skill once on a held-out split. On ContextBench's KV Store task at 32K, starting from *no* instruction, the evolved skill raises held-out accuracy from **38.3% to 74.2%** — a **35.9-point** gain ($74.2 - 38.3 = 35.9$) — at lower compute. *(reasoned, matching the reported "35.9 points.")*

**Reinforcement learning** internalizes the strategy into the weights. Because context edits change the input across turns, they use stepwise GRPO: sample a group of full trajectories, compute the usual outcome advantage, and assign it to every model call in the trajectory. The new piece is a **success-gated efficiency advantage**. The problem it solves: rewarding "fewer tokens" or "more edits" directly invites reward hacking — the model learns to make pointless edits that throw away information or wreck prefix reuse. So the efficiency signal only ranks *among already-successful* trajectories. For successful trajectory $i$ in group $g$ with prefix-reuse FLOPs $c_i$ and group mean $\bar{c}_g$:

$$A_i^{\mathrm{eff}} = \operatorname{clip}\!\left(\frac{\bar{c}_g - c_i}{\bar{c}_g},\, -1,\, 1\right)$$

and the total advantage is $A_i = A_i^{\mathrm{out}} + w_{\mathrm{eff}} A_i^{\mathrm{eff}}$. Failed trajectories, and groups with fewer than two successes, get $A_i^{\mathrm{eff}} = 0$. Efficiency never trades away correctness; it only breaks ties between correct runs by cost.

Training Qwen3.5-9B on OpenResearcher and evaluating on held-out BrowseComp-Plus:

| | Acc. (%) | PFLOPs / Q |
|---|---|---|
| Summary harness | 34.7 → 42.1 | 4.01 → 2.19 |
| CLM | 28.8 → 42.5 | 1.52 → 1.34 |

Before RL, the small CLM sits at **28.8%**, six points behind the summary harness — a 9B model is not a great context manager out of the box. After RL it gains **13.7 points to 42.5%**, the abstract's **47.6%** relative improvement ($42.5/28.8 = 1.476$), while its per-question cost falls from 1.52 to 1.34 PFLOPs — **12% fewer** ($(1.52-1.34)/1.52 = 0.118$). That jump from 28.8 to 42.5 is the strategy being learned in the weights. *(reasoned, from the reported table.)* It ends **0.4 points above** the trained summary harness while spending **38.8% fewer** FLOPs than it ($(2.19-1.34)/2.19 = 0.388$). *(reasoned.)* The policy that a human would have hard-coded into a harness is instead discovered and folded into the weights.

## What arbitrary edits cost to serve

There is a tax, and the paper is honest about it. Modern servers make long contexts affordable through **prefix caching**: the KV cache is content-addressed, so a turn that only *appends* reuses everything computed before and prefills just the new tokens. (If you want the why, [self-attention](/articles/how-transformers-attention-works) ties each token's key and value to every token before it and, through rotary position encodings, to its absolute position.) An edit *in the middle* breaks that: every token after the first mismatch has to be re-prefilled, including text that did not change.

To make the cost legible, the authors define **prefix-reuse FLOPs** and work it out for Qwen3.6-27B, a hybrid model with 48 of its 64 layers using linear attention. Per-token projection cost is $C_{\mathrm{token}} = 48.70 \times 10^{9}$ FLOPs and each query-key pair costs $C_{\mathrm{attn}} = 3.93 \times 10^{5}$. For a turn with prompt length $P$, generated length $G$, and a reusable prefix of $R$ tokens, the cost is

$$F = C_{\mathrm{token}}(P - R + G) + C_{\mathrm{attn}}\!\left[\tfrac{1}{2}(P^2 - R^2) + GP + \tfrac{1}{2}G^2\right]$$

I re-implemented this formula from the constants above and confirmed it reproduces the paper's three worked points exactly. Drag the edit position:

<ReprefillCost />

For an illustrative turn with $P = 20{,}000$ and $G = 500$: appending (reusing 18,000 tokens) costs $1.41 \times 10^{14}$ FLOPs, and prefix caching has avoided **87%** of what a cold turn would need. An edit in the middle that leaves 10,000 reusable tokens raises it to $5.74 \times 10^{14}$. An edit at the very start reuses nothing and costs $10.81 \times 10^{14}$ — **7.7 times** the append-only turn ($10.81/1.41 = 7.67$). *(measured; the formula reproduces the paper's reported $1.41$, $5.74$, $10.81$ and 7.7×.)* The important thing is that the paper reports *all* its efficiency numbers under this metric, under standard serving — so the "fewer FLOPs" wins above already pay the edit tax. CLMs come out ahead not by avoiding re-prefills but by keeping the context so small that each re-prefill is cheap.

<Figure src="https://ai.thesatyajit.com/articles/context-language-models/fig2.png" alt="Diagram. Original context is three blocks A, B, C. An edit replaces B with B-prime. Standard serving reuses the cache for A, prefills B-prime, and re-prefills all of C. Suffix Cache Reuse reuses A, prefills B-prime, and reuses the cache for C as well." caption="Suffix Cache Reuse relocates the surviving suffix's cached states past the edit instead of re-prefilling them (Context Language Models paper, Figure 4)." />

Then they attack the tax directly with **Suffix Cache Reuse (SCR)**, a patch to SGLang. When an edit replaces $B$ with $B'$ in a context $[A\,B\,C]$, standard serving re-prefills $B'$ *and* the surviving suffix $C$. SCR instead reuses $C$'s cached keys and values: it diffs the new prompt against the old, finds the surviving spans, **re-rotates their rotary position encodings** to the new positions, and splices them in after $B'$. Only $B'$ is prefilled. It relocates up to $K$ spans per edit, largest first; the released config defaults to **K=6** (`KVREUSE_MAX_BLOCKS=6`), which I confirmed in `suffix_cache_reuse/config.py`. *(measured, from the code.)* The reused states are slightly stale — computed under the pre-edit context — so SCR *approximates* re-prefilling, and the cap bounds how much approximation any one edit introduces.

The result: SCR matches standard SGLang at **65.0% of its prefix-reuse FLOPs** on BrowseComp-Plus — a **35% reduction** in server-side compute at matched performance. *(reported.)* A nice side effect: reasoning models often strip earlier turns' reasoning blocks before the next turn, which also forces a re-prefill, and SCR fixes that too even when the model never edits its own context. Of the 7.8% of prompt tokens SCR reuses beyond prefix-cache hits, 5.3 points come from reasoning-stripping and only 2.5 from edits — so most of the benefit applies to ordinary reasoning-model serving. *(reported.)*

## Where it is still a promise

The thesis is strong and the receipts are mostly real, but a few things deserve flags.

- **I verified arithmetic and code structure, not the runs.** Every benchmark number here is **reported** — the training and 24-hour evaluations are not reproducible in a sandbox, and I did not re-run them. What I independently checked is the serving-cost formula (it reproduces the paper's figures from its own constants), the context-as-a-file mechanism (`ContextEnv` really does mirror, edit, and sync), and the SCR defaults. Treat the benchmark wins as the authors' measurements, not mine.
- **ContextBench is not released yet.** The repo's README lists it as "coming soon," so the one benchmark built specifically to isolate context management — and the source of the 35.9-point gain — cannot be independently run today. The harness, the RL code, and SCR *are* released.
- **The license is research-only.** The code is **CC BY-NC 4.0**, not a permissive license. Fine for study; not something you drop into a commercial agent.
- **Context-length awareness is a real limitation.** To decide *when* to edit, a model has to know how full its context is, and the paper's own appendix shows existing LMs are poor at this — they guess bucketed values rather than reading their true length, and need external hints to calibrate. A CLM that cannot feel its own budget leans back on the harness to tell it — exactly the dependency the paper is trying to remove.

None of that undercuts the core finding, which I think is right: context management was the last big piece of the agent loop still living entirely in hand-written scaffolding, and this is a credible argument — with numbers under an honest cost metric — that it does not have to. The harness gets eaten; the context goes first.
