# Recursive Self-Rewrite: teach the model what the harness did

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/recursive-self-rewrite
> date: 2026-10-06
> tags: explainer, agents, llm, fine-tuning, self-improvement, benchmarks, qwen, open-weights

The site keeps returning to one uncomfortable fact about agents: the loop around the model decides
a lot of the score. [Agent harnesses](/articles/agent-harness) made the case in general,
[the harness effect](/articles/harness-effect) put a token bill on it, and
[the harness is the generalizer](/articles/harness-compositional-generalization) argued the
loop can carry generalization the weights don't. Hugging Face's [multi-harness RL guide](/articles/multi-harness-rl) makes the
same point from the training side: the same weights score 62.1% under Mini-SWE-Agent and 33.2%
under Claude Code.

Most of the work that follows from that fact improves the harness and leaves the model alone.
[RRSI](/articles/rrsi-harness-evolution) and [Recursive Harness Self-Improvement](/articles/recursive-harness-self-improvement)
evolve the loop; [Dream-RSI](/articles/dream-rsi) evolves a scheduler. **Recursive Self-Rewrite**
(RSR; Li et al., [arXiv 2610.02826](https://arxiv.org/abs/2610.02826), 2 Oct 2026) goes the other way. It uses
several harnesses as *discovery tools*, collects whatever they help the model solve, and then
moves that skill into the weights, so the model can do it again under a plain loop with none of
that help. The released checkpoint is [IntelligenceLab/RSR-27B](https://huggingface.co/IntelligenceLab/RSR-27B)
on Hugging Face, Apache-2.0.

Every number below is labelled. **Reported** is the paper's or the model card's figure, not re-run.
**Measured** is something I read off a file myself. **Reasoned** is my arithmetic on the other two.
I did not train or evaluate anything.

## The problem: a success belongs to the model and the loop together

A harness decides how the model sees observations, calls tools, tracks progress, checks its work,
recovers and stops. Change it and you change which tasks get solved. RSR uses three, all driving the
same Qwen-3.8-27B on terminal tasks:

- **Terminus 2**, the general loop from the Harbor framework. The model issues shell commands,
  reads the output, and decides when it is done. This is also the *target* harness: the one the
  model will be deployed and evaluated under.
- **StateM**, which organizes a run around persistent workflow states, phase-local context and
  checked transitions: write a contract, pass a check, move to the next phase.
- **RSRT** (Recursive Self-Reflect Terminus), the authors' modification of Terminus 2. When the
  model claims it is done, the harness runs the task verifier. On a fail it reports only "failed",
  never the verifier's internals, and tells the model to review its trajectory, debug and continue,
  up to 180 minutes of wall clock per task (reported).

<Figure
  src="https://ai.thesatyajit.com/articles/recursive-self-rewrite/fig3.png"
  alt="Three panels. Terminus 2: model and terminal in an act-observe loop, the model decides when to finish. StateM: workflow state plus context feeds the model, a transition check either repairs or passes back to the state. RSRT: model works in terminal, finishes, task verifier passes to stop or fails into reflect plus debug, which loops back to the model."
  caption="The three discovery harnesses: general interaction, state-guided execution, and outcome-guided repair. The same base model runs under all three (RSR paper, Figure 3)."
/>

The three are complementary, and the paper measures it well. On a pool of 2,929 terminal tasks
(2,509 self-built SWR tasks across about 50 domains plus 420 harder RST tasks), the model ran
14,598 discovery rollouts and passed 2,001 of them, 13.7% (reported). Those passes cover 759
distinct tasks. The best single harness, Terminus 2, covers 565; the union adds 194, a 34.3%
relative increase (reported; 194 / 565 = 34.3%, reasoned). Of the 759, 288 are solved by exactly one
harness (129 Terminus 2, 91 RSRT, 68 StateM), 163 by two, 308 by all three (reported; the regions
in Figure 4b sum to 759, measured).

<Figure
  src="https://ai.thesatyajit.com/articles/recursive-self-rewrite/fig4.png"
  alt="Left: grouped bar chart of the share of tasks solved by Terminus 2, RSRT, StateM and their union on RST (420), SWR (2,509) and pooled (2,929); the union is highest everywhere, 30.7, 25.1 and 25.9 percent. Right: an upset plot of solve regions on the pooled set: 308 tasks solved by all three, 129 by Terminus 2 only, 91 by RSRT only, 78 by Terminus 2 and RSRT, 68 by StateM only, 50 by Terminus 2 and StateM, 35 by RSRT and StateM."
  caption="Coverage of each harness and their union, and the exclusive and overlapping solve regions. Union 759, best single 565, additional 194 (RSR paper, Figure 4)."
/>

The raw pool has unequal rollout counts (5,405 Terminus 2, 3,777 RSRT, 5,416 StateM; reported,
RSRT had more sandbox errors), so the authors also check a balanced subset of 1,247 tasks with
equal attempts per harness. There RSRT is the strongest single harness at 285 tasks, the union
reaches 352, and each harness still has exclusive solves: 30, 56 and 21 (reported). The
complementarity is not an artefact of sampling.

So a model with three loops can solve a quarter more tasks than a model with one. The catch is what
those extra successes look like. A StateM trajectory is full of StateM: phase resets, transition
checks, commands like `statem goto implement`. An RSRT trajectory contains the verifier's
"failed, keep going" messages and runs past the point where the model itself would have stopped:
RSRT averages 2.20 completion claims per passing trajectory against 1.11 for StateM, and 4.5% of
its passes come after a rejected completion (reported). Train on those as-is and you teach the
model to expect a controller that will not be there at inference time.

The paper has a direct measurement of what that costs. **Direct SFT**, fine-tuning on the passing
source rollouts with the harness-specific insertion phrases stripped out, *lowers* Terminal-Bench 2
pass@3 from 57.0% to 53.4% and its per-run mean from 51.7% to 43.8% (reported). Inspecting three
trajectories from that checkpoint, the authors saw a recurring dead end: the model loops on the same
action pattern without progress. RSRT produces many trajectories longer than 100 steps, and the
paper's reading is that a 27B model imitating those cannot sustain them without the harness that
kept prodding it (reported; three inspected trajectories is an anecdote, and the paper says as much).

## The rewrite: re-solve it, don't edit it

The key design decision is that RSR never edits a trajectory. Stripping control messages from a
StateM run would leave a log with holes where the controller used to push. Instead, the same base
model plays three roles and produces a *new* trajectory from a fresh start.

<Figure
  src="https://ai.thesatyajit.com/articles/recursive-self-rewrite/fig2.png"
  alt="Three-stage pipeline. Discover: tasks fan out to harness 1, 2, 3, producing more difficult solutions. Passing traces go to Rewrite: a planner writes K runbooks per trace with milestones and checks; a critic does rule checks and model checks for solution, verifier and harness leakage, sending leaks back to the planner to resample with the reason; the approved private guide goes to the executor under the general harness, acting in a fresh sandbox, M times per runbook. Rewritten traces go to Verify: task verifier with hidden checks, fail discards, pass goes to a leak audit for underived values, then a filter, producing verified trajectories."
  caption="Discover with several harnesses, rewrite with planner, critic and executor (one model in all three roles), then verify and audit before anything becomes training data (RSR paper, Figure 2)."
/>

**Compact.** The source trajectory is cut down to the task instruction, the model's actions and the
environment's observations. Harness control messages are removed. This compacted log is input for the
planner only; it never becomes training data.

**Plan.** The planner reads the compacted success and writes a *runbook*: the required end state, key
milestones, useful checks, recovery strategies and common pitfalls. It may describe the task, the
interfaces and how to validate. It must not hand over the finished deliverable. It samples K = 4
runbook candidates per source trajectory (reported).

**Criticize.** Every candidate goes through deterministic checks first (schema validation, removal of
known artefacts, rejection of references to tools the general harness doesn't have), then a
model-based critic. The critic sees only the public task instruction and the candidate runbook, and
decides whether the runbook teaches a procedure or leaks something the executor shouldn't get: the
verifier's checks, the solution, harness-specific machinery. A rejected runbook goes back to the
planner with the critic's reason, and the planner regenerates. That loop, capped at N iterations in
the paper's Figure 1, is the "recursive" in the name.

**Execute.** For each accepted runbook, the executor solves the task again, in a fresh sandbox, under
plain Terminus 2, M = 4 times at temperature 0.7 (reported). The runbook is private guidance: it is in
the executor's context while it generates, but it is never written into the trajectory. So the
executor has to react to what this sandbox actually prints, rather than replay the old run.

**Filter.** A rewrite has to pass the task's verifier. Then the model is used once more to flag any
value that appears in the trajectory but can't be derived from the task or the environment, a sign
that an answer travelled through the runbook, and those rewrites are thrown away. What is kept is
the public interaction history: instruction, observations, the model's own responses. The runbook
and the critic conversation are dropped.

The widget walks one of the paper's own cases through those stages: an OpenFOAM reconstruction task
first solved under StateM in 38 turns. The commands marked "fig 6" are the ones the paper prints for
the source and its passing rewrite; the connecting lines are mine, there to show what each stage
removes.

<RewriteStepper />

Two things make the result trustworthy as training data, at least by construction. Correctness is
re-established rather than inherited: every kept rewrite passed the verifier in its own sandbox, so
it is a real success under the target harness, not a cleaned-up transcript of someone else's. And the
leak defence is layered: the critic screens the runbook before execution, and the underived-value
audit screens the trajectory after. Neither is a proof. The critic and the auditor are the same 27B
model that is being trained, and the paper reports no false-negative rate for either.

<Figure
  src="https://ai.thesatyajit.com/articles/recursive-self-rewrite/fig1.png"
  alt="Concept illustration. Left, success experiences: one model under harness A, B and C, each with a passing trace. Middle, recursive planning: the planner writes a runbook with goal, steps, check and avoid lines, and a struck-out solution line flagged as a leak; the critic rejects drafts 1 and 2 back to the planner with the reason, at most N times, and accepts draft 3. Right, re-learn experience: runbooks 1, 2 and 3 each drive several make-test-fix attempts, most passing and one discarded; one success becomes many, up to K times M."
  caption="The planner-critic loop is where the recursion lives: a runbook goes back with the critic's reason until it is clean or N attempts run out, and one success becomes up to K × M fresh attempts (RSR paper, Figure 1)."
/>

## Is it recursive in the self-improvement sense?

No, and the paper is explicit about it. Its related-work section says RSR "does not optimize a
single harness or assume an autonomous recursive cycle." There is one round: the base model discovers,
the base model rewrites, the base model is fine-tuned once. The improved model is not sent back to
rediscover and rewrite again. So the brief question "does the improved model rewrite again?" has a
plain answer: not in this paper. A second round is the obvious experiment and nothing here measures it.
Compare [Dream-RSI](/articles/dream-rsi), where what recurses across rounds is the search code,
and [RRSI](/articles/rrsi-harness-evolution), where it is the harness; RSR's only loop is inside the
data generator.

The training is equally plain: supervised fine-tuning, no RL. The RSR model is trained on the
verified rewrites plus the 766 direct passing Terminus 2 trajectories, which need no rewriting because
they are already in the target harness (reported). The paper gives no learning rate, epoch count,
sequence length or compute budget, so the recipe cannot be reproduced from the text alone.

## How many trajectories come out

The rewrite stage targets the StateM and RSRT successes, 636 + 599 = 1,235 source trajectories
(reported counts; sum reasoned). With K = 4 runbooks and M = 4 executions each, the ceiling is 16
attempts per source, about 19,760 in total (reasoned). The paper reports 12,893 rewrite rollouts, so
roughly a third of runbook slots must have been dropped by the critic or lost to sandbox errors
(reasoned). Of those 12,893, 86% pass, giving 11,094 training trajectories (reported; 11,094 / 12,893
= 86.0%, reasoned).

One count does not reconcile. The paper says the rewrites cover 975 unique tasks, but discovery
solved only 759 tasks in total, and the StateM and RSRT sources cover at most 630 of them (Table 1's
RSRT ∪ StateM row). A rewrite cannot cover a task no source solved. Either 975 counts something
else, or the rewrite stage ran on sources beyond those described. I can't resolve it from the paper.

The two case studies give a feel for yield. For a Markdown inline-parsing task from RSRT, 7 of 12
rewrites pass (58.3%), and the median passing rewrite takes 32 turns against the source's 64
(reported). For the OpenFOAM task from StateM, 7 of 8 pass (87.5%), median 29 turns against 38, and
the rewrites keep the source's analyse-implement-repair rhythm without any StateM commands (reported).
Rewrites guided by the same runbook share more commands than rewrites from different runbooks (19.1
vs 12.1 Jaccard on Markdown, 12.6 vs 10.4 on OpenFOAM; reported), which is the evidence that the
runbook steers the procedure without dictating the commands.

<Figure
  src="https://ai.thesatyajit.com/articles/recursive-self-rewrite/fig6.png"
  alt="Two case studies. Case 1, Markdown inline parsing from an RSRT source, 7 of 12 pass, 64 to 32 median turns; the runbook reads find a compatible parser, match the output summary, probe the table parameter. The source reaches the prebuilt tree-sitter parser at turn 51; the passing rewrite uses it at turn 9; a failing rewrite hits xxd command not found. Case 2, OpenFOAM demo outputs from a StateM source, 7 of 8 pass, 38 to 29 median turns; the source runs tb_receipt.py validate-contract and statem goto implement; the passing rewrite inspects reconstruct.py and rebuilds; a failing rewrite hits a NameError and a KeyError."
  caption="A source run, a passing rewrite and a failing rewrite guided by the same runbook, for two tasks. The rewrite under Terminus 2 has no StateM commands in it (RSR paper, Figure 6)."
/>

The model card shows three more tasks the same way, with wall-clock time. All three sources ran
under continue-until-timeout and hit the limit, and the passing rewrites are less than half as long:
64 → 34 steps and 2.0 h → 22 min on a tree-sitter task, 59 → 27 steps on CMake, 56 → 23 on OpenFOAM
(reported). The card calls it a 2-hour limit; the paper says RSRT allows up to 180 minutes. These
are hand-picked successes, each with 4 of 4 replays passing; the paper's own 7-of-12 case is the more
representative yield.

<Figure
  src="https://ai.thesatyajit.com/articles/recursive-self-rewrite/card-trajectories.png"
  alt="Horizontal bars comparing an expert trajectory and a passing rewrite for three SWR tasks, coloured by command kind (explore, edit, run, install, wait, other). tree_sitter_markdown_inline_05: expert 64 steps, 153 commands, 2.0 h; rewrite 34 steps, 78 commands, 22 min. cmake_generated_header_config_06: expert 59 steps, 94 commands, 2.0 h; rewrite 27 steps, 58 commands, 21 min. openfoam_scalar_transport_boundary_probe_02: expert 56 steps, 72 commands, 1.9 h; rewrite 23 steps, 51 commands, 28 min."
  caption="Expert (base model under continue-until-timeout) versus a passing rewrite (base model under plain Terminus 2, runbook kept private) on three sample SWR tasks; selected examples, not averages (IntelligenceLab/RSR-27B model card)."
/>

## Results

All three models are evaluated under plain Terminus 2, three runs per benchmark. Pass@3 counts a task
if any of the three runs passes; the mean averages the three per-run pass rates.

| Benchmark (tasks) | Metric | Base | Direct SFT | RSR |
|---|---|---|---|---|
| Terminal-Bench 2 (89) | pass@3 | 57.0% | 53.4% | **74.2%** |
| Terminal-Bench 3 (74) | pass@3 | 0.0% | 5.4% | **9.5%** |
| Terminal-Bench 4 (66) | pass@3 | 1.5% | 4.5% | **9.1%** |
| TBH, authors' (100) | pass@3 | 39.0% | 56.0% | **63.0%** |
| SWR100, authors' (100) | pass@3 | 3.0% | 3.0% | **6.0%** |
| Terminal-Bench 2 | mean of 3 | 51.7% | 43.8% | **70.1%** |
| TBH | mean of 3 | 31.0% | 42.5% | **54.0%** |
| LHTB, authors' (46) | process reward | 0.21 | 0.25 | **0.29** |

All reported (paper, Table 5). RSR beats Direct SFT on every row: by 20.8, 4.1, 4.6, 7.0 and 3.0 points
of pass@3 and by 26.3, 5.4, 2.9, 11.5 and 2.0 points of per-run mean on the five pass-rate benchmarks
(reported, and the differences check out, reasoned). On Long-Horizon Terminal-Bench, all three
complete 0 of 46 tasks; the process-reward gain is partial progress, not solved tasks.

<ResultsBoard />

## What the numbers do and don't show

**The headline holds on the public benchmark.** Terminal-Bench 2 is not the authors' own, and
+17.2 points of pass@3 and +18.4 of per-run mean over the base model (reasoned from Table 5) is a large
move for one round of SFT on self-generated data. The contrast with Direct SFT, which went *down*, is
the paper's best evidence that the rewrite step matters and not just the extra successes.

**It is not a controlled data comparison.** Direct SFT trains on the raw passing rollouts; the paper
doesn't give its count, but the pool it describes holds 2,001. RSR trains on 11,094 rewrites plus 766 direct trajectories, 11,860 in all (reported counts,
sum reasoned), about 5.9 times as many (reasoned). The paper has no size-matched ablation, and no
"RSR without the critic" or "rewrite only Terminus 2" arm. So I can't separate "cleaner trajectories"
from "six times more trajectories" in the 20.8-point gap. Some of both, presumably.

**Two of the benchmarks are the authors' own, and one is close to the training data.** TBH comes from
the same RST pipeline (Li et al., 2026a) that supplied 420 of the training tasks. The paper doesn't
describe any overlap check between the TBH tasks and the training pool. The TBH gains may be fine,
but they carry less weight than the Terminal-Bench 2 gains.

**The pass@3 counts are not exact fractions.** Table 5 prints "57.0% (51/89)" and notes that the counts
are back-computed from the percentages. But 51/89 is 57.3%, and no integer over 89 gives 57.0% or
53.4% (reasoned). The percentages come from somewhere other than a plain count over 89 tasks, perhaps
tasks that failed to run, and the paper doesn't say. It doesn't change the story; it is a reason to read
the decimals loosely.

**The harness alone gets you most of the way, on their own benchmark.** Table 6 runs the *untrained*
base model once under each discovery harness. On Terminal-Bench 2 that gives 51.7% (Terminus 2),
55.1% (RSRT), 58.4% (StateM). On TBH it gives 33.0%, 66.0% and 46.0% (reported). The untrained model
under RSRT, at 66.0% in a single run, beats the trained RSR model under Terminus 2 on TBH: 63.0% pass@3,
54.0% per run. On Terminal-Bench 2 the trained model wins clearly, 70.1% per run against StateM's 58.4%.
That is the trade the paper is making explicit: RSR is worth it when you want one cheap, general loop at
deployment, not the best possible score on a task family where a heavier loop already works. (Table 6
also prints the TB3 union as "4.3 (3)"; 3/74 is 4.1%, reasoned.)

## The released model

[IntelligenceLab/RSR-27B](https://huggingface.co/IntelligenceLab/RSR-27B) is a `Qwen3_5ForConditionalGeneration`
checkpoint, 64 layers, hidden size 5,120, with 55.56 GB of safetensors across 16 main shards plus two
named `model-missing-from-origin-*` (measured from `config.json`, the file listing and
`model.safetensors.index.json`). Those two hold the vision tower (333 tensors) and the multi-token
prediction head (15 tensors), measured. My reading is that the fine-tune touched the language model and
the rest was copied back in from the base checkpoint; the card doesn't say. 55.56 GB at 2 bytes per
parameter is about 27.8B parameters (reasoned, assuming BF16). The card has no evaluation table, only
the three sample trajectories above and a link to a trajectory viewer.

## What I take from it

The mechanism is simple enough to copy, and the insight underneath it is the useful part: **a success
under a heavy harness is a hint, not a demonstration.** You don't train on it; you have the model solve
the task again, under the loop it will actually run in, with the hint kept out of the log. Correctness
comes back from the verifier, not from the original run.

What it leaves open is everything a second round would answer. Does RSR-27B under the three
harnesses discover tasks the base model couldn't, and does rewriting those help again, or does the
gain flatten after one pass? The name promises recursion at that level, and this paper only does it
inside the runbook loop. It also leaves the boundary where a heavier loop still beats a trained model,
which Table 6 suggests is real, unmapped.

Related on the site: [Scaling agentic RL](/articles/scaling-agentic-rl) and
[MiMo-V2.6's RL environments](/articles/mimo-rl-environments) on where verifiable terminal tasks come
from, and [Qwen3.8-Max](/articles/qwen3-8-max) for the family RSR's base model belongs to.
