2026-10-02 · 16 min · explainer · agents · llm · systems · benchmarks · evaluation
Agent harnesses made the case that the loop around a model — the prompts, the tool interfaces, the control flow, the memory and the context policy — sets what an agent can do about as much as the weights do. The harness effect showed the same loop decides the token bill. Both treat the harness as something an engineer tunes by hand, reading failed traces one at a time. The obvious next move is to let a model read those traces and rewrite the harness itself. Recursive Harness Self-Improvement did it with a pairwise comparison against the previous version; a dozen other methods do it by proposing component-wise edits and keeping whatever raises a benchmark score.
That last clause is the problem. Keeping whatever raises a benchmark score, round after round, on the same finite set of tasks, is a recipe for memorizing the set. RRSI (Xia, Han, Wang, … Pfister, Lee — Google Research, Stanford, WashU, UNC, arXiv 2609.24972, 21 Sep 2026) is a direct answer to that failure mode. It keeps the harness fully editable but regularizes how the search moves, borrowing the logic of L0, L1 and L2 regularization and pointing it at the search trajectory instead of a weight vector. The code is on GitHub under Apache-2.0, and the numbers below are the paper's unless I say otherwise; I checked them against that repo's configs and against Harvey LAB, the legal-work benchmark RRSI evolves on.
Improve the model, or improve the harness
An agent is a pair : a frozen backbone policy and a harness . The harness is everything around the weights — the system and task prompts, the control flow that decides when the agent plans, acts, reflects or stops, the tool descriptions, the skill and memory files, and the context management that decides what the policy sees at each step. On a task the agent produces a trajectory and a deliverable, scored by a verifier as . Over a task set you measure a score and a policy-token cost .
Two knobs can raise . You can change — pretrain or fine-tune the weights — which is expensive and, for most teams using a frozen frontier model, not available. Or you can change , which is a program edit: a prompt rewrite, a new tool, a different stopping rule. Harness evolution is the automated version of the second knob. At round it runs the current harness on an evolve set, summarizes the trajectories into feedback, asks a proposer LLM for candidate edits, scores the candidates on the same evolve set, and keeps the best:
This is recursive self-improvement at the system level: the system's own output picks the next version of the system. It is also, in the paper's framing, "adaptive empirical optimization over an unusually expressive search space" — and that is exactly the setup where overfitting lives.
Why naive self-improvement overfits
The evolve set is reused adaptively. The candidates proposed at round depend on scores measured on the same tasks in earlier rounds, so every evaluation is another look at the same data, and the search slowly shapes itself to that data's idiosyncrasies. Three failure modes fall out of that, and the paper names all three:
- Benchmark-specific fitting. The proposer can encode the evolve suite's task names, entity names or answers directly into a prompt or a skill file. That lifts the evolve score and transfers nothing.
- Noise chasing. Each is an average over a handful of stochastic runs. Pick the largest measured score across many candidates and rounds, and you will often be picking the luckiest noise draw, not the best mechanism.
- Complexity accumulation. Adding machinery — more retries, more sub-agents, more context — can nudge the evolve score up while spending tokens and improving nothing fundamental.
All three widen the gap between the evolve score and held-out performance. A harness that improves on the split it is scored on and regresses everywhere else has memorized the split. That is the thing to prevent.
A regularization view of the search
Here is the move that makes RRSI more than a list of heuristics. Let be the set of harnesses reachable from by arbitrary source edits. Most "safe" approaches would shrink — allow only prompt edits, say. RRSI leaves wide open (prompts, control flow, config, context management, tools, skills, memory and sub-agents may all change) and instead constrains the trajectory through it. The constraints are lifted straight from classical regularization:
- An edit budget that caps how many independent edits one candidate bundles is an -style cardinality constraint on the update.
- Complexity-aware acceptance, which refuses cost growth that is not paid for by measured gain, is Ridge / -style shrinkage of the harness's resource footprint.
- Structural pruning, which deletes components that stop earning their keep, is Lasso / -style sparsification of the retained harness.
The analogy is qualitative — there is no penalized objective being minimized — but it organizes the whole design, and it is the reason the regularizers compose instead of fighting each other.

The seven constraints
Proposal side: control how capacity is spent
Annealed update sparsity (). An unconstrained proposer can bundle a dozen unrelated changes into one candidate; any score move is then impossible to attribute to a mechanism. RRSI caps the number of independently attributable edits per candidate with a cosine schedule that decays from to over a -round run:
Early rounds may combine several coordinated changes to discover a mechanism; later rounds get
sparse and attributable. In the repo's configs the coding instance runs
over rounds, and the engineering instance over (measured, from
domains/*/rrsi.json).
Evidence-aware credit assignment. Constraining update size only helps if the search remembers what earlier updates established. RRSI records, per evaluated candidate, the component it touched, the hypothesis it tested, the diff, the resulting score and cost change, and whether it was accepted. The proposer conditions on that whole history, so a falsified hypothesis is not drawn again and a successful mechanism keeps explicit credit. This is the part that treats the evolve set like the adaptive-data-analysis problem it is: repeatedly re-testing something you already falsified spends looks at the data without buying evidence.
Structured exploration. The same history reveals when the search has collapsed onto one edit family — rewriting prompts over and over while the structural machinery sits untouched. RRSI calls the run stalled when its progress over a window stays inside the noise band , and during a stall it reserves part of the budget for component types never exercised in the run. It is entropy regularization for a code search: push capacity toward the unexplored without restricting what the harness may contain.
Selection side: control what becomes permanent
A candidate has to clear several non-compensatory gates before its score can justify replacing the incumbent — a big win on one axis cannot buy a failure on another.
Leakage screening. Before a candidate is scored, a critic reads its diff and rejects anything that encodes task names, entity names, task-specific values, answers, or inert machinery. Screening before evaluation is the subtle part: a leaking candidate never receives the inflated evolve score that would make it attractive to later rounds. Generic prompt or tool-description improvements pass; a lookup table keyed by the suite's task names does not.
Stability-aware acceptance. Before evolution, RRSI repeatedly evaluates the unchanged base harness and estimates an empirical noise band . A candidate must then clear a floor set at the best score seen so far minus that band:
The floor stops the search from walking downhill through a run of regressions each too small to
tell from noise. The bands are calibrated per domain and are small: for coding
(three passes out of trials), for the Harvey LAB workspace (about 60
criteria out of ~14,100), and for engineering (five passes out of ) —
all measured from rrsi.json. A gain under the band is treated as a coin flip, not a win.
Complexity-aware acceptance (Ridge / ). For a candidate whose gain clears the band, added cost has to be justified by the gain:
where is the relative policy-token change and the score change. is the cost tolerated for a negligible gain; is how much more cost each extra point of gain buys. Policy tokens are the proxy for the harness's resource footprint, and this rule shrinks that footprint without forcing any one component out.
Structural pruning (Lasso / ). The budget sparsifies each update; pruning sparsifies the retained harness. RRSI tracks whether a recently exercised component produced a positive gain over a pruning window, and reports the dead ones to the proposer as deletion targets. A mechanism has to keep earning its place rather than persist because score-only evolution has no reason to remove it.
The interactive below runs the selection side on a handful of candidate edits. The numbers are made up to show the decision; the real bands are the ones above. The leakage critic kills the hard-coded lookup before it is ever scored; the floor swallows the cosmetic prompt rewrite, and swallows more as you widen ; the cost rule rejects the five-sub-agent edit whose small gain does not pay for its tokens. Only the edits that actually transfer survive.
Does it transfer?
The whole point is held-out performance, so every number below is measured against the unevolved base harness in the same window, with Claude Opus 4.8 frozen as the policy. "Evolve" is the split the harness was searched on; every other row never entered selection (all reported, from the paper and the repo README).
| Domain | Benchmark | Role | RRSI | Δ | |
|---|---|---|---|---|---|
| Coding | Terminal-Bench 2.1 | evolve | 74.2 | 80.2 | +6.0 |
| Coding | SWE-bench Verified | OOD | 82.0 | 83.8 | +1.8 |
| Workspace | Harvey LAB | evolve | 89.4 | 90.5 | +1.1 |
| Workspace | Harvey LAB | ID held-out | 86.9 | 89.2 | +2.3 |
| Workspace | JobBench | OOD | 36.0 | 40.7 | +4.7 |
| Workspace | GDPval | OOD | 48.8 | 52.3 | +3.5 |
| Workspace | APEX-Agents | OOD | 34.2 | 37.9 | +3.7 |
| Engineering | EngDesign | evolve | 50.0 | 54.9 | +4.9 |
| Engineering | Frontier-Eng | OOD | 17.7 | 22.0 | +4.3 |

The claim to check is that the evolve gain survives the move off the evolve split. On coding, Terminal-Bench 2.1 goes 74.2 → 80.2 (+6.0 on the split it was searched on), and SWE-bench Verified — repository-level bug fixing, never scored during the search — still gains 1.8 points. No held-out split regresses anywhere, which is precisely the failure a memorizing harness produces and which Figure 1 shows the baselines producing. Table 1 in the paper makes the trade explicit: RRSI posts the smallest evolve-set gain of any evolved harness and the only out-of-distribution average that clears by more than a point (43.6 vs 39.7, reported). The strongest baseline on the evolve split adds 0.9 points out of distribution; two baselines finish below where they started. Giving up evolve-set score to keep transfer is the deal the regularizers are designed to make.
The paper's abstract headline — "up to 14.1 points on the split it evolves against" — is not in the Opus table; it comes from the policy-robustness run. Swap the frozen policy to Gemini 3.5 Flash and the same coding instance goes 64.6 → 78.7 on Terminal-Bench 2.1 (+14.1) and transfers +2.2 to SWE-bench Verified (reported). A weaker base leaves more room, so the headline number rides on the weaker policy; the Opus gains are smaller because Opus starts closer to the ceiling. Either way the harness improves a benchmark it was never scored on, under two unrelated backbones — and, in a further run, the Gemini-evolved harness still helps Gemini 3.1 Flash Lite, a model that never took part in the search (11.2 → 14.6). A harness is a program, not a set of weights; a mechanism that only helps the policy it was searched against is an artifact, and these do not read as artifacts.
The cheaper-harness claim
The other headline is "30% fewer policy tokens than the unregularized evolution." The workspace ablation (Table 2, reported) is where to check it:
| Variant | LAB evolve | LAB held-out | OOD avg | Tokens/trial (M) |
|---|---|---|---|---|
| (no evolution) | 89.4 | 86.9 | 39.7 | 1.56 |
| Unregularized evolution | 92.8 | 88.9 | 40.3 | 3.80 |
| w/o proposal regularizers | 90.7 | 88.8 | 41.9 | 2.69 |
| w/o acceptance regularizers | 91.5 | 88.7 | 41.0 | 3.59 |
| RRSI | 90.5 | 89.2 | 43.6 | 2.42 |
Unregularized evolution posts the highest evolve score of any arm (92.8) and the lowest transfer (40.3, within a point of doing nothing), at 3.80M tokens per trial. RRSI runs at 2.42M. That is a 36% reduction (reasoned, ), so the paper's "30%" is if anything conservative for this instance; I read it as an across-domain headline rather than this one number. The ablation also settles which half does what: dropping the acceptance constraints lifts the evolve score and sinks transfer while cost rises by half, and dropping the proposal constraints costs almost nothing on the evolve split but 1.7 points out of distribution. Steering where the search looks matters even when nothing is rejected. No evolved harness is as cheap as (1.56M) — evolution does buy part of its gain with test-time compute, and the budget decides how much.
Why Harvey LAB is the right target
Self-improvement is only as trustworthy as the benchmark you measure it on, and most agent
benchmarks are the wrong shape for it: a single pass/fail per task, a public leaderboard to climb,
and grading loose enough that writing the way a judge rewards beats doing the work.
Harvey LAB (Harvey AI,
MIT-licensed, lab-core v1.2.0) is built the other way, and three of its choices are exactly what a
self-improving system has to be held to.
It is real work, graded the way the work is checked. LAB is a filesystem-first benchmark of
legal tasks: the agent reads a synthetic matter — a data room, a set of contracts — and writes
deliverables, which an LLM judge grades against a rubric defined inline in each task. The launch
post (May 2026) described "more than 1,200 agent tasks across 24 legal practice areas … over 75,000
expert-written rubric criteria." The repo today is larger: I counted 2,010 task.json files across
27 top-level practice-area directories (measured; the README badge reads tasks-2010), and the eval
docs put it at ~114,000 rubric criteria. The task count in the original thread — 1,671 — is a real
number this set passed through on its way up.
Grading is all-pass. A task scores 1.0 only if every binary criterion passes, else 0.0. No
partial credit at the task level and none within a criterion. Harvey's own framing: a diligence memo
that catches eight of ten risks is not 80% useful, it is materially incomplete, and the missing
issue could change the deal. For a system that edits itself against a score, all-pass grading is a
feature — it removes the easy points a graded mean would hand to cosmetic edits, and it is a harsh,
honest gradient. The standard profile runs two judges, claude-sonnet-4-6 and gpt-5.5, and the
aggregate all-pass requires both to all-pass the task — one judge agreeing is not enough.
It launched with no leaderboard, on purpose. Harvey says it is "intentionally launching LAB without a leaderboard because we expect the dataset to evolve over time." A benchmark with no ranking to climb and a task set that keeps changing is a poor thing to overfit and a good thing to measure transfer against — which is precisely why RRSI uses a fixed 120-task evolve split and a pristine 40-task held-out split carved from a pinned LAB commit, and reports both. The one external LAB score the RRSI paper cites, 89.4 → 90.5 on evolve and 86.9 → 89.2 held-out, is the transfer story in miniature: a small evolve gain that survives the move to tasks the search never saw.
What it costs, and what it doesn't fix
The honest limits. RRSI still relies on a finite evolve set and a fistful of hyperparameters — the budget bounds, the cost coefficients, the noise band — so it inherits a dependence on feedback quality and search budget. Every evolved harness costs more than ; the win is transfer per token against other evolution methods, not free capability. And the whole study freezes the weights. It says nothing about model updates and harness updates sharing a loop, which is where the scaffolding-gets-eaten warning lives: a mechanism a harness installs today can be absorbed by a stronger model tomorrow, repricing the whole search.
What RRSI gets right is the framing. Autonomous research runs, engineering swarms and agent planners all share the structure it formalizes: a system proposing edits to itself under an empirical fitness signal. The lesson is that recursive self-improvement at the system level is a generalization problem, not only a search problem — and what you control is not what the harness may contain but how repeated, noisy feedback becomes permanent state. Leave the edit space open; regularize the path through it; measure on work graded the way the work is checked.