~/satyajit

RRSI: regularize the harness search, not the harness

mdjsonmcp

2026-10-02 · 16 min · explainer · agents · llm · systems · benchmarks · evaluation

Agent harnesses made the case that the loop around a model — the prompts, the tool interfaces, the control flow, the memory and the context policy — sets what an agent can do about as much as the weights do. The harness effect showed the same loop decides the token bill. Both treat the harness as something an engineer tunes by hand, reading failed traces one at a time. The obvious next move is to let a model read those traces and rewrite the harness itself. Recursive Harness Self-Improvement did it with a pairwise comparison against the previous version; a dozen other methods do it by proposing component-wise edits and keeping whatever raises a benchmark score.

That last clause is the problem. Keeping whatever raises a benchmark score, round after round, on the same finite set of tasks, is a recipe for memorizing the set. RRSI (Xia, Han, Wang, … Pfister, Lee — Google Research, Stanford, WashU, UNC, arXiv 2609.24972, 21 Sep 2026) is a direct answer to that failure mode. It keeps the harness fully editable but regularizes how the search moves, borrowing the logic of L0, L1 and L2 regularization and pointing it at the search trajectory instead of a weight vector. The code is on GitHub under Apache-2.0, and the numbers below are the paper's unless I say otherwise; I checked them against that repo's configs and against Harvey LAB, the legal-work benchmark RRSI evolves on.

Improve the model, or improve the harness

An agent is a pair A=(π,H)A = (\pi, H): a frozen backbone policy π\pi and a harness HH. The harness is everything around the weights — the system and task prompts, the control flow that decides when the agent plans, acts, reflects or stops, the tool descriptions, the skill and memory files, and the context management that decides what the policy sees at each step. On a task xx the agent produces a trajectory τ\tau and a deliverable, scored by a verifier as r(x,τ)∈[0,1]r(x,\tau) \in [0,1]. Over a task set you measure a score S^(H)\hat{S}(H) and a policy-token cost C^(H)\hat{C}(H).

Two knobs can raise S^\hat{S}. You can change π\pi — pretrain or fine-tune the weights — which is expensive and, for most teams using a frozen frontier model, not available. Or you can change HH, which is a program edit: a prompt rewrite, a new tool, a different stopping rule. Harness evolution is the automated version of the second knob. At round tt it runs the current harness HtH_t on an evolve set, summarizes the trajectories into feedback, asks a proposer LLM for candidate edits, scores the candidates on the same evolve set, and keeps the best:

Ht+1=arg max⁡H′∈Ht∪{Ht}S^(H′;Devolve).H_{t+1} = \operatorname*{arg\,max}_{H' \in \mathcal{H}_t \cup \{H_t\}} \hat{S}(H'; \mathcal{D}_{\mathrm{evolve}}).

This is recursive self-improvement at the system level: the system's own output picks the next version of the system. It is also, in the paper's framing, "adaptive empirical optimization over an unusually expressive search space" — and that is exactly the setup where overfitting lives.

Why naive self-improvement overfits

The evolve set is reused adaptively. The candidates proposed at round tt depend on scores measured on the same tasks in earlier rounds, so every evaluation is another look at the same data, and the search slowly shapes itself to that data's idiosyncrasies. Three failure modes fall out of that, and the paper names all three:

All three widen the gap between the evolve score and held-out performance. A harness that improves on the split it is scored on and regresses everywhere else has memorized the split. That is the thing to prevent.

Here is the move that makes RRSI more than a list of heuristics. Let Ω(H)\Omega(H) be the set of harnesses reachable from HH by arbitrary source edits. Most "safe" approaches would shrink Ω(H)\Omega(H) — allow only prompt edits, say. RRSI leaves Ω(H)\Omega(H) wide open (prompts, control flow, config, context management, tools, skills, memory and sub-agents may all change) and instead constrains the trajectory through it. The constraints are lifted straight from classical regularization:

The analogy is qualitative — there is no penalized objective being minimized — but it organizes the whole design, and it is the reason the regularizers compose instead of fighting each other.

RRSI overview: a central regularized transition rule flanked by proposal-side constraints (annealed update sparsity, evidence-aware credit assignment, structured exploration) on the left and selection-side constraints (leakage screening, a noise-adjusted performance floor, complexity-aware acceptance, structural pruning) on the right.
RRSI regularizes both halves of the loop: proposal-side constraints control how search capacity is spent, selection-side constraints control which measured gains become permanent state (RRSI, Figure 2).

The seven constraints

Proposal side: control how capacity is spent

Annealed update sparsity (L0L_0). An unconstrained proposer can bundle a dozen unrelated changes into one candidate; any score move is then impossible to attribute to a mechanism. RRSI caps the number of independently attributable edits per candidate with a cosine schedule that decays from bmax⁡b_{\max} to bmin⁡b_{\min} over a TT-round run:

bt=⌈bmin⁡+(bmax⁡−bmin⁡)⋅12(1+cos⁡(πt/T))⌉.b_t = \left\lceil b_{\min} + (b_{\max} - b_{\min}) \cdot \tfrac{1}{2}\bigl(1 + \cos(\pi t / T)\bigr) \right\rceil.

Early rounds may combine several coordinated changes to discover a mechanism; later rounds get sparse and attributable. In the repo's configs the coding instance runs bmax⁡=4→bmin⁡=1b_{\max}=4 \to b_{\min}=1 over T=20T=20 rounds, and the engineering instance 4→14 \to 1 over T=40T=40 (measured, from domains/*/rrsi.json).

Evidence-aware credit assignment. Constraining update size only helps if the search remembers what earlier updates established. RRSI records, per evaluated candidate, the component it touched, the hypothesis it tested, the diff, the resulting score and cost change, and whether it was accepted. The proposer conditions on that whole history, so a falsified hypothesis is not drawn again and a successful mechanism keeps explicit credit. This is the part that treats the evolve set like the adaptive-data-analysis problem it is: repeatedly re-testing something you already falsified spends looks at the data without buying evidence.

Structured exploration. The same history reveals when the search has collapsed onto one edit family — rewriting prompts over and over while the structural machinery sits untouched. RRSI calls the run stalled when its progress over a window stays inside the noise band δ\delta, and during a stall it reserves part of the budget for component types never exercised in the run. It is entropy regularization for a code search: push capacity toward the unexplored without restricting what the harness may contain.

Selection side: control what becomes permanent

A candidate has to clear several non-compensatory gates before its score can justify replacing the incumbent — a big win on one axis cannot buy a failure on another.

Leakage screening. Before a candidate is scored, a critic reads its diff and rejects anything that encodes task names, entity names, task-specific values, answers, or inert machinery. Screening before evaluation is the subtle part: a leaking candidate never receives the inflated evolve score that would make it attractive to later rounds. Generic prompt or tool-description improvements pass; a lookup table keyed by the suite's task names does not.

Stability-aware acceptance. Before evolution, RRSI repeatedly evaluates the unchanged base harness and estimates an empirical noise band δ\delta. A candidate must then clear a floor set at the best score seen so far minus that band:

S^(H′)≥S⋆−δ.\hat{S}(H') \ge S^{\star} - \delta.

The floor stops the search from walking downhill through a run of regressions each too small to tell from noise. The bands are calibrated per domain and are small: δ=0.017\delta = 0.017 for coding (three passes out of 89×2=17889 \times 2 = 178 trials), 0.0040.004 for the Harvey LAB workspace (about 60 criteria out of ~14,100), and 0.0200.020 for engineering (five passes out of 61×4=24461 \times 4 = 244) — all measured from rrsi.json. A gain under the band is treated as a coin flip, not a win.

Complexity-aware acceptance (Ridge / L2L_2). For a candidate whose gain clears the band, added cost has to be justified by the gain:

ΔC≤β0+β1 ΔS,\Delta C \le \beta_0 + \beta_1\,\Delta S,

where ΔC\Delta C is the relative policy-token change and ΔS\Delta S the score change. β0\beta_0 is the cost tolerated for a negligible gain; β1\beta_1 is how much more cost each extra point of gain buys. Policy tokens are the proxy for the harness's resource footprint, and this rule shrinks that footprint without forcing any one component out.

Structural pruning (Lasso / L1L_1). The budget sparsifies each update; pruning sparsifies the retained harness. RRSI tracks whether a recently exercised component produced a positive gain over a pruning window, and reports the dead ones to the proposer as deletion targets. A mechanism has to keep earning its place rather than persist because score-only evolution has no reason to remove it.

The interactive below runs the selection side on a handful of candidate edits. The numbers are made up to show the decision; the real bands are the ones above. The leakage critic kills the hard-coded lookup before it is ever scored; the floor swallows the cosmetic prompt rewrite, and swallows more as you widen δ\delta; the cost rule rejects the five-sub-agent edit whose small gain does not pay for its tokens. Only the edits that actually transfer survive.

Selection, gate by gate
illustrative
proposed edit → control flow
Re-read the file it edited and re-run the check before it stops.
evolve gain +2.8 ptscost +18.0%
1. Leakage critic
No benchmark-specific logic in the diff.
2. Noise floor
Gain 2.8 > band 2.0.
3. Cost rule
+18.0% cost, paid for by the gain.
Accepted → becomes the next incumbent
Held-out gain that actually transfers: +2.2 pts.

Does it transfer?

The whole point is held-out performance, so every number below is measured against the unevolved base harness H0H_0 in the same window, with Claude Opus 4.8 frozen as the policy. "Evolve" is the split the harness was searched on; every other row never entered selection (all reported, from the paper and the repo README).

DomainBenchmarkRoleH0H_0RRSIΔ
CodingTerminal-Bench 2.1evolve74.280.2+6.0
CodingSWE-bench VerifiedOOD82.083.8+1.8
WorkspaceHarvey LABevolve89.490.5+1.1
WorkspaceHarvey LABID held-out86.989.2+2.3
WorkspaceJobBenchOOD36.040.7+4.7
WorkspaceGDPvalOOD48.852.3+3.5
WorkspaceAPEX-AgentsOOD34.237.9+3.7
EngineeringEngDesignevolve50.054.9+4.9
EngineeringFrontier-EngOOD17.722.0+4.3
Grouped bar chart across nine benchmark columns, each with three bars: the unevolved harness H0 in grey, RRSI on the split it is scored on in light blue, and RRSI on a split it never sees in dark blue. Every column rises from grey to the RRSI bar, and no held-out column regresses.
The headline result across all three domains. Grey is the unevolved harness; the gains on splits the search never scored on are the ones that matter (RRSI, Figure 3).

The claim to check is that the evolve gain survives the move off the evolve split. On coding, Terminal-Bench 2.1 goes 74.2 → 80.2 (+6.0 on the split it was searched on), and SWE-bench Verified — repository-level bug fixing, never scored during the search — still gains 1.8 points. No held-out split regresses anywhere, which is precisely the failure a memorizing harness produces and which Figure 1 shows the baselines producing. Table 1 in the paper makes the trade explicit: RRSI posts the smallest evolve-set gain of any evolved harness and the only out-of-distribution average that clears H0H_0 by more than a point (43.6 vs 39.7, reported). The strongest baseline on the evolve split adds 0.9 points out of distribution; two baselines finish below where they started. Giving up evolve-set score to keep transfer is the deal the regularizers are designed to make.

The paper's abstract headline — "up to 14.1 points on the split it evolves against" — is not in the Opus table; it comes from the policy-robustness run. Swap the frozen policy to Gemini 3.5 Flash and the same coding instance goes 64.6 → 78.7 on Terminal-Bench 2.1 (+14.1) and transfers +2.2 to SWE-bench Verified (reported). A weaker base leaves more room, so the headline number rides on the weaker policy; the Opus gains are smaller because Opus starts closer to the ceiling. Either way the harness improves a benchmark it was never scored on, under two unrelated backbones — and, in a further run, the Gemini-evolved harness still helps Gemini 3.1 Flash Lite, a model that never took part in the search (11.2 → 14.6). A harness is a program, not a set of weights; a mechanism that only helps the policy it was searched against is an artifact, and these do not read as artifacts.

The cheaper-harness claim

The other headline is "30% fewer policy tokens than the unregularized evolution." The workspace ablation (Table 2, reported) is where to check it:

VariantLAB evolveLAB held-outOOD avgTokens/trial (M)
H0H_0 (no evolution)89.486.939.71.56
Unregularized evolution92.888.940.33.80
w/o proposal regularizers90.788.841.92.69
w/o acceptance regularizers91.588.741.03.59
RRSI90.589.243.62.42

Unregularized evolution posts the highest evolve score of any arm (92.8) and the lowest transfer (40.3, within a point of doing nothing), at 3.80M tokens per trial. RRSI runs at 2.42M. That is a 36% reduction (reasoned, (3.80−2.42)/3.80(3.80-2.42)/3.80), so the paper's "30%" is if anything conservative for this instance; I read it as an across-domain headline rather than this one number. The ablation also settles which half does what: dropping the acceptance constraints lifts the evolve score and sinks transfer while cost rises by half, and dropping the proposal constraints costs almost nothing on the evolve split but 1.7 points out of distribution. Steering where the search looks matters even when nothing is rejected. No evolved harness is as cheap as H0H_0 (1.56M) — evolution does buy part of its gain with test-time compute, and the budget decides how much.

Why Harvey LAB is the right target

Self-improvement is only as trustworthy as the benchmark you measure it on, and most agent benchmarks are the wrong shape for it: a single pass/fail per task, a public leaderboard to climb, and grading loose enough that writing the way a judge rewards beats doing the work. Harvey LAB (Harvey AI, MIT-licensed, lab-core v1.2.0) is built the other way, and three of its choices are exactly what a self-improving system has to be held to.

It is real work, graded the way the work is checked. LAB is a filesystem-first benchmark of legal tasks: the agent reads a synthetic matter — a data room, a set of contracts — and writes deliverables, which an LLM judge grades against a rubric defined inline in each task. The launch post (May 2026) described "more than 1,200 agent tasks across 24 legal practice areas … over 75,000 expert-written rubric criteria." The repo today is larger: I counted 2,010 task.json files across 27 top-level practice-area directories (measured; the README badge reads tasks-2010), and the eval docs put it at ~114,000 rubric criteria. The task count in the original thread — 1,671 — is a real number this set passed through on its way up.

Grading is all-pass. A task scores 1.0 only if every binary criterion passes, else 0.0. No partial credit at the task level and none within a criterion. Harvey's own framing: a diligence memo that catches eight of ten risks is not 80% useful, it is materially incomplete, and the missing issue could change the deal. For a system that edits itself against a score, all-pass grading is a feature — it removes the easy points a graded mean would hand to cosmetic edits, and it is a harsh, honest gradient. The standard profile runs two judges, claude-sonnet-4-6 and gpt-5.5, and the aggregate all-pass requires both to all-pass the task — one judge agreeing is not enough.

It launched with no leaderboard, on purpose. Harvey says it is "intentionally launching LAB without a leaderboard because we expect the dataset to evolve over time." A benchmark with no ranking to climb and a task set that keeps changing is a poor thing to overfit and a good thing to measure transfer against — which is precisely why RRSI uses a fixed 120-task evolve split and a pristine 40-task held-out split carved from a pinned LAB commit, and reports both. The one external LAB score the RRSI paper cites, 89.4 → 90.5 on evolve and 86.9 → 89.2 held-out, is the transfer story in miniature: a small evolve gain that survives the move to tasks the search never saw.

What it costs, and what it doesn't fix

The honest limits. RRSI still relies on a finite evolve set and a fistful of hyperparameters — the budget bounds, the cost coefficients, the noise band — so it inherits a dependence on feedback quality and search budget. Every evolved harness costs more than H0H_0; the win is transfer per token against other evolution methods, not free capability. And the whole study freezes the weights. It says nothing about model updates and harness updates sharing a loop, which is where the scaffolding-gets-eaten warning lives: a mechanism a harness installs today can be absorbed by a stronger model tomorrow, repricing the whole search.

What RRSI gets right is the framing. Autonomous research runs, engineering swarms and agent planners all share the structure it formalizes: a system proposing edits to itself under an empirical fitness signal. The lesson is that recursive self-improvement at the system level is a generalization problem, not only a search problem — and what you control is not what the harness may contain but how repeated, noisy feedback becomes permanent state. Leave the edit space open; regularize the path through it; measure on work graded the way the work is checked.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "RRSI: regularize the harness search, not the harness", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026rrsiharnessevolution,
  author = {Satyajit Ghana},
  title  = {RRSI: regularize the harness search, not the harness},
  url    = {https://ai.thesatyajit.com/articles/rrsi-harness-evolution},
  year   = {2026}
}
share