# RRSI: regularize the harness search, not the harness

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/rrsi-harness-evolution
> date: 2026-10-02
> tags: explainer, agents, llm, systems, benchmarks, evaluation

[Agent harnesses](/articles/agent-harness) made the case that the loop around a model — the
prompts, the tool interfaces, the control flow, the memory and the context policy — sets what
an agent can do about as much as the weights do. [The harness effect](/articles/harness-effect)
showed the same loop decides the token bill. Both treat the harness as something an engineer
tunes by hand, reading failed traces one at a time. The obvious next move is to let a model read
those traces and rewrite the harness itself. [Recursive Harness Self-Improvement](/articles/recursive-harness-self-improvement)
did it with a pairwise comparison against the previous version; a dozen other methods do it by
proposing component-wise edits and keeping whatever raises a benchmark score.

That last clause is the problem. Keeping whatever raises a benchmark score, round after round, on
the same finite set of tasks, is a recipe for memorizing the set.
[RRSI](https://arxiv.org/abs/2609.24972) (Xia, Han, Wang, … Pfister, Lee — Google Research, Stanford,
WashU, UNC, arXiv 2609.24972, 21 Sep 2026) is a direct answer to that failure mode. It keeps the harness
fully editable but regularizes *how the search moves*, borrowing the logic of L0, L1 and L2
regularization and pointing it at the search trajectory instead of a weight vector. The
[code is on GitHub](https://github.com/google-research/rrsi) under Apache-2.0, and the numbers
below are the paper's unless I say otherwise; I checked them against that repo's configs and against
[Harvey LAB](https://github.com/harveyai/harvey-labs), the legal-work benchmark RRSI evolves on.

## Improve the model, or improve the harness

An agent is a pair $A = (\pi, H)$: a frozen backbone policy $\pi$ and a harness $H$. The harness
is everything around the weights — the system and task prompts, the control flow that decides when
the agent plans, acts, reflects or stops, the tool descriptions, the skill and memory files, and
the context management that decides what the policy sees at each step. On a task $x$ the agent
produces a trajectory $\tau$ and a deliverable, scored by a verifier as $r(x,\tau) \in [0,1]$. Over
a task set you measure a score $\hat{S}(H)$ and a policy-token cost $\hat{C}(H)$.

Two knobs can raise $\hat{S}$. You can change $\pi$ — pretrain or fine-tune the weights — which is
expensive and, for most teams using a frozen frontier model, not available. Or you can change $H$,
which is a program edit: a prompt rewrite, a new tool, a different stopping rule. Harness evolution
is the automated version of the second knob. At round $t$ it runs the current harness $H_t$ on an
evolve set, summarizes the trajectories into feedback, asks a proposer LLM for candidate edits,
scores the candidates on the same evolve set, and keeps the best:

$$
H_{t+1} = \operatorname*{arg\,max}_{H' \in \mathcal{H}_t \cup \{H_t\}} \hat{S}(H'; \mathcal{D}_{\mathrm{evolve}}).
$$

This is recursive self-improvement at the system level: the system's own output picks the next
version of the system. It is also, in the paper's framing, "adaptive empirical optimization over an
unusually expressive search space" — and that is exactly the setup where overfitting lives.

## Why naive self-improvement overfits

The evolve set is reused *adaptively*. The candidates proposed at round $t$ depend on scores
measured on the same tasks in earlier rounds, so every evaluation is another look at the same data,
and the search slowly shapes itself to that data's idiosyncrasies. Three failure modes fall out of
that, and the paper names all three:

- **Benchmark-specific fitting.** The proposer can encode the evolve suite's task names, entity
  names or answers directly into a prompt or a skill file. That lifts the evolve score and transfers
  nothing.
- **Noise chasing.** Each $\hat{S}$ is an average over a handful of stochastic runs. Pick the
  largest measured score across many candidates and rounds, and you will often be picking the
  luckiest noise draw, not the best mechanism.
- **Complexity accumulation.** Adding machinery — more retries, more sub-agents, more context — can
  nudge the evolve score up while spending tokens and improving nothing fundamental.

All three widen the gap between the evolve score and held-out performance. A harness that improves
on the split it is scored on and regresses everywhere else has memorized the split. That is the
thing to prevent.

## A regularization view of the search

Here is the move that makes RRSI more than a list of heuristics. Let $\Omega(H)$ be the set of
harnesses reachable from $H$ by arbitrary source edits. Most "safe" approaches would shrink
$\Omega(H)$ — allow only prompt edits, say. RRSI leaves $\Omega(H)$ wide open (prompts, control
flow, config, context management, tools, skills, memory and sub-agents may all change) and instead
constrains the *trajectory* through it. The constraints are lifted straight from classical
regularization:

- An **edit budget** that caps how many independent edits one candidate bundles is an $L_0$-style
  cardinality constraint on the update.
- **Complexity-aware acceptance**, which refuses cost growth that is not paid for by measured gain,
  is Ridge / $L_2$-style shrinkage of the harness's resource footprint.
- **Structural pruning**, which deletes components that stop earning their keep, is Lasso /
  $L_1$-style sparsification of the retained harness.

The analogy is qualitative — there is no penalized objective being minimized — but it organizes the
whole design, and it is the reason the regularizers compose instead of fighting each other.

<Figure
  src="https://ai.thesatyajit.com/articles/rrsi-harness-evolution/fig1.png"
  alt="RRSI overview: a central regularized transition rule flanked by proposal-side constraints (annealed update sparsity, evidence-aware credit assignment, structured exploration) on the left and selection-side constraints (leakage screening, a noise-adjusted performance floor, complexity-aware acceptance, structural pruning) on the right."
  caption="RRSI regularizes both halves of the loop: proposal-side constraints control how search capacity is spent, selection-side constraints control which measured gains become permanent state (RRSI, Figure 2)."
/>

## The seven constraints

### Proposal side: control how capacity is spent

**Annealed update sparsity ($L_0$).** An unconstrained proposer can bundle a dozen unrelated
changes into one candidate; any score move is then impossible to attribute to a mechanism. RRSI
caps the number of independently attributable edits per candidate with a cosine schedule that
decays from $b_{\max}$ to $b_{\min}$ over a $T$-round run:

$$
b_t = \left\lceil b_{\min} + (b_{\max} - b_{\min}) \cdot \tfrac{1}{2}\bigl(1 + \cos(\pi t / T)\bigr) \right\rceil.
$$

Early rounds may combine several coordinated changes to discover a mechanism; later rounds get
sparse and attributable. In the repo's configs the coding instance runs $b_{\max}=4 \to b_{\min}=1$
over $T=20$ rounds, and the engineering instance $4 \to 1$ over $T=40$ (measured, from
`domains/*/rrsi.json`).

**Evidence-aware credit assignment.** Constraining update size only helps if the search remembers
what earlier updates established. RRSI records, per evaluated candidate, the component it touched,
the hypothesis it tested, the diff, the resulting score and cost change, and whether it was
accepted. The proposer conditions on that whole history, so a falsified hypothesis is not drawn
again and a successful mechanism keeps explicit credit. This is the part that treats the evolve set
like the adaptive-data-analysis problem it is: repeatedly re-testing something you already
falsified spends looks at the data without buying evidence.

**Structured exploration.** The same history reveals when the search has collapsed onto one edit
family — rewriting prompts over and over while the structural machinery sits untouched. RRSI calls
the run *stalled* when its progress over a window stays inside the noise band $\delta$, and during a
stall it reserves part of the budget for component types never exercised in the run. It is entropy
regularization for a code search: push capacity toward the unexplored without restricting what the
harness may contain.

### Selection side: control what becomes permanent

A candidate has to clear several non-compensatory gates before its score can justify replacing the
incumbent — a big win on one axis cannot buy a failure on another.

**Leakage screening.** *Before* a candidate is scored, a critic reads its diff and rejects anything
that encodes task names, entity names, task-specific values, answers, or inert machinery. Screening
before evaluation is the subtle part: a leaking candidate never receives the inflated evolve score
that would make it attractive to later rounds. Generic prompt or tool-description improvements pass;
a lookup table keyed by the suite's task names does not.

**Stability-aware acceptance.** Before evolution, RRSI repeatedly evaluates the unchanged base
harness and estimates an empirical noise band $\delta$. A candidate must then clear a floor set at
the best score seen so far minus that band:

$$
\hat{S}(H') \ge S^{\star} - \delta.
$$

The floor stops the search from walking downhill through a run of regressions each too small to
tell from noise. The bands are calibrated per domain and are small: $\delta = 0.017$ for coding
(three passes out of $89 \times 2 = 178$ trials), $0.004$ for the Harvey LAB workspace (about 60
criteria out of ~14,100), and $0.020$ for engineering (five passes out of $61 \times 4 = 244$) —
all measured from `rrsi.json`. A gain under the band is treated as a coin flip, not a win.

**Complexity-aware acceptance (Ridge / $L_2$).** For a candidate whose gain clears the band, added
cost has to be justified by the gain:

$$
\Delta C \le \beta_0 + \beta_1\,\Delta S,
$$

where $\Delta C$ is the relative policy-token change and $\Delta S$ the score change. $\beta_0$ is
the cost tolerated for a negligible gain; $\beta_1$ is how much more cost each extra point of gain
buys. Policy tokens are the proxy for the harness's resource footprint, and this rule shrinks that
footprint without forcing any one component out.

**Structural pruning (Lasso / $L_1$).** The budget sparsifies each update; pruning sparsifies the
*retained* harness. RRSI tracks whether a recently exercised component produced a positive gain
over a pruning window, and reports the dead ones to the proposer as deletion targets. A mechanism
has to keep earning its place rather than persist because score-only evolution has no reason to
remove it.

The interactive below runs the selection side on a handful of candidate edits. The numbers are made
up to show the decision; the real bands are the ones above. The leakage critic kills the hard-coded
lookup before it is ever scored; the floor swallows the cosmetic prompt rewrite, and swallows more
as you widen $\delta$; the cost rule rejects the five-sub-agent edit whose small gain does not pay
for its tokens. Only the edits that actually transfer survive.

<EvolveGate />

## Does it transfer?

The whole point is held-out performance, so every number below is measured against the *unevolved*
base harness $H_0$ in the same window, with Claude Opus 4.8 frozen as the policy. "Evolve" is the
split the harness was searched on; every other row never entered selection (all reported, from the
paper and the repo README).

| Domain | Benchmark | Role | $H_0$ | RRSI | Δ |
|---|---|---|---:|---:|---:|
| Coding | Terminal-Bench 2.1 | evolve | 74.2 | 80.2 | +6.0 |
| Coding | SWE-bench Verified | OOD | 82.0 | 83.8 | +1.8 |
| Workspace | Harvey LAB | evolve | 89.4 | 90.5 | +1.1 |
| Workspace | Harvey LAB | ID held-out | 86.9 | 89.2 | +2.3 |
| Workspace | JobBench | OOD | 36.0 | 40.7 | +4.7 |
| Workspace | GDPval | OOD | 48.8 | 52.3 | +3.5 |
| Workspace | APEX-Agents | OOD | 34.2 | 37.9 | +3.7 |
| Engineering | EngDesign | evolve | 50.0 | 54.9 | +4.9 |
| Engineering | Frontier-Eng | OOD | 17.7 | 22.0 | +4.3 |

<Figure
  src="https://ai.thesatyajit.com/articles/rrsi-harness-evolution/fig2.png"
  alt="Grouped bar chart across nine benchmark columns, each with three bars: the unevolved harness H0 in grey, RRSI on the split it is scored on in light blue, and RRSI on a split it never sees in dark blue. Every column rises from grey to the RRSI bar, and no held-out column regresses."
  caption="The headline result across all three domains. Grey is the unevolved harness; the gains on splits the search never scored on are the ones that matter (RRSI, Figure 3)."
/>

The claim to check is that the evolve gain survives the move off the evolve split. On coding,
Terminal-Bench 2.1 goes 74.2 → 80.2 (+6.0 on the split it was searched on), and SWE-bench Verified
— repository-level bug fixing, never scored during the search — still gains 1.8 points. No held-out
split regresses anywhere, which is precisely the failure a memorizing harness produces and which
Figure 1 shows the baselines producing. Table 1 in the paper makes the trade explicit: RRSI posts
the *smallest* evolve-set gain of any evolved harness and the only out-of-distribution average that
clears $H_0$ by more than a point (43.6 vs 39.7, reported). The strongest baseline on the evolve
split adds 0.9 points out of distribution; two baselines finish below where they started. Giving up
evolve-set score to keep transfer is the deal the regularizers are designed to make.

The paper's abstract headline — "up to 14.1 points on the split it evolves against" — is not in the
Opus table; it comes from the policy-robustness run. Swap the frozen policy to Gemini 3.5 Flash and
the same coding instance goes 64.6 → 78.7 on Terminal-Bench 2.1 (+14.1) and transfers +2.2 to
SWE-bench Verified (reported). A weaker base leaves more room, so the headline number rides on the
weaker policy; the Opus gains are smaller because Opus starts closer to the ceiling. Either way the
harness improves a benchmark it was never scored on, under two unrelated backbones — and, in a
further run, the Gemini-evolved harness still helps Gemini 3.1 Flash Lite, a model that never took
part in the search (11.2 → 14.6). A harness is a program, not a set of weights; a mechanism that
only helps the policy it was searched against is an artifact, and these do not read as artifacts.

## The cheaper-harness claim

The other headline is "30% fewer policy tokens than the unregularized evolution." The workspace
ablation (Table 2, reported) is where to check it:

| Variant | LAB evolve | LAB held-out | OOD avg | Tokens/trial (M) |
|---|---:|---:|---:|---:|
| $H_0$ (no evolution) | 89.4 | 86.9 | 39.7 | 1.56 |
| Unregularized evolution | 92.8 | 88.9 | 40.3 | 3.80 |
| w/o proposal regularizers | 90.7 | 88.8 | 41.9 | 2.69 |
| w/o acceptance regularizers | 91.5 | 88.7 | 41.0 | 3.59 |
| RRSI | 90.5 | 89.2 | 43.6 | 2.42 |

Unregularized evolution posts the highest evolve score of any arm (92.8) and the lowest transfer
(40.3, within a point of doing nothing), at 3.80M tokens per trial. RRSI runs at 2.42M. That is a
36% reduction (reasoned, $(3.80-2.42)/3.80$), so the paper's "30%" is if anything conservative for
this instance; I read it as an across-domain headline rather than this one number. The ablation also
settles which half does what: dropping the acceptance constraints lifts the evolve score and sinks
transfer while cost rises by half, and dropping the proposal constraints costs almost nothing on the
evolve split but 1.7 points out of distribution. Steering where the search looks matters even when
nothing is rejected. No evolved harness is as cheap as $H_0$ (1.56M) — evolution does buy part of
its gain with test-time compute, and the budget decides how much.

<Callout type="note">
One honest caveat on the Harvey LAB numbers specifically. RRSI's workspace instance scores LAB as
the *fraction of rubric criteria passed*, under a single Gemini 3.5 Flash judge (`judge_model` in
`domains/workspace/rrsi.json`). That is not LAB's own headline metric, which is the all-pass rate
under a two-judge panel. So 89.4 → 90.5 is a criterion-pass fraction, not an all-pass rate — useful
as an evolve signal, not comparable to a published LAB all-pass score.
</Callout>

## Why Harvey LAB is the right target

Self-improvement is only as trustworthy as the benchmark you measure it on, and most agent
benchmarks are the wrong shape for it: a single pass/fail per task, a public leaderboard to climb,
and grading loose enough that writing the way a judge rewards beats doing the work.
[Harvey LAB](https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark) (Harvey AI,
MIT-licensed, `lab-core` v1.2.0) is built the other way, and three of its choices are exactly what a
self-improving system has to be held to.

**It is real work, graded the way the work is checked.** LAB is a filesystem-first benchmark of
legal tasks: the agent reads a synthetic matter — a data room, a set of contracts — and writes
deliverables, which an LLM judge grades against a rubric defined inline in each task. The launch
post (May 2026) described "more than 1,200 agent tasks across 24 legal practice areas … over 75,000
expert-written rubric criteria." The repo today is larger: I counted 2,010 `task.json` files across
27 top-level practice-area directories (measured; the README badge reads `tasks-2010`), and the eval
docs put it at ~114,000 rubric criteria. The task count in the original thread — 1,671 — is a real
number this set passed through on its way up.

**Grading is all-pass.** A task scores 1.0 only if *every* binary criterion passes, else 0.0. No
partial credit at the task level and none within a criterion. Harvey's own framing: a diligence memo
that catches eight of ten risks is not 80% useful, it is materially incomplete, and the missing
issue could change the deal. For a system that edits itself against a score, all-pass grading is a
feature — it removes the easy points a graded mean would hand to cosmetic edits, and it is a harsh,
honest gradient. The standard profile runs two judges, `claude-sonnet-4-6` and `gpt-5.5`, and the
aggregate all-pass requires *both* to all-pass the task — one judge agreeing is not enough.

**It launched with no leaderboard, on purpose.** Harvey says it is "intentionally launching LAB
without a leaderboard because we expect the dataset to evolve over time." A benchmark with no
ranking to climb and a task set that keeps changing is a poor thing to overfit and a good thing to
measure transfer against — which is precisely why RRSI uses a fixed 120-task evolve split and a
pristine 40-task held-out split carved from a pinned LAB commit, and reports both. The one external
LAB score the RRSI paper cites, 89.4 → 90.5 on evolve and 86.9 → 89.2 held-out, is the transfer
story in miniature: a small evolve gain that survives the move to tasks the search never saw.

## What it costs, and what it doesn't fix

The honest limits. RRSI still relies on a finite evolve set and a fistful of hyperparameters — the
budget bounds, the cost coefficients, the noise band — so it inherits a dependence on feedback
quality and search budget. Every evolved harness costs more than $H_0$; the win is transfer per
token against other evolution methods, not free capability. And the whole study freezes the weights.
It says nothing about model updates and harness updates sharing a loop, which is where
[the scaffolding-gets-eaten](/articles/scaffolding-gets-eaten) warning lives: a mechanism a harness
installs today can be absorbed by a stronger model tomorrow, repricing the whole search.

What RRSI gets right is the framing. [Autonomous research runs](/articles/nanogpt-speedrun-frontier),
[engineering swarms](/articles/jev-engineering-swarms) and [agent planners](/articles/prime-agent)
all share the structure it formalizes: a system proposing edits to itself under an empirical fitness
signal. The lesson is that recursive self-improvement at the system level is a *generalization*
problem, not only a search problem — and what you control is not what the harness may contain but
how repeated, noisy feedback becomes permanent state. Leave the edit space open; regularize the path
through it; measure on work graded the way the work is checked.
