~/satyajit

Recursive Self-Rewrite: teach the model what the harness did

mdjsonmcp

2026-10-06 · 17 min · explainer · agents · llm · fine-tuning · self-improvement · benchmarks · qwen · open-weights

The site keeps returning to one uncomfortable fact about agents: the loop around the model decides a lot of the score. Agent harnesses made the case in general, the harness effect put a token bill on it, and the harness is the generalizer argued the loop can carry generalization the weights don't. Hugging Face's multi-harness RL guide makes the same point from the training side: the same weights score 62.1% under Mini-SWE-Agent and 33.2% under Claude Code.

Most of the work that follows from that fact improves the harness and leaves the model alone. RRSI and Recursive Harness Self-Improvement evolve the loop; Dream-RSI evolves a scheduler. Recursive Self-Rewrite (RSR; Li et al., arXiv 2610.02826, 2 Oct 2026) goes the other way. It uses several harnesses as discovery tools, collects whatever they help the model solve, and then moves that skill into the weights, so the model can do it again under a plain loop with none of that help. The released checkpoint is IntelligenceLab/RSR-27B on Hugging Face, Apache-2.0.

Every number below is labelled. Reported is the paper's or the model card's figure, not re-run. Measured is something I read off a file myself. Reasoned is my arithmetic on the other two. I did not train or evaluate anything.

The problem: a success belongs to the model and the loop together

A harness decides how the model sees observations, calls tools, tracks progress, checks its work, recovers and stops. Change it and you change which tasks get solved. RSR uses three, all driving the same Qwen-3.8-27B on terminal tasks:

Three panels. Terminus 2: model and terminal in an act-observe loop, the model decides when to finish. StateM: workflow state plus context feeds the model, a transition check either repairs or passes back to the state. RSRT: model works in terminal, finishes, task verifier passes to stop or fails into reflect plus debug, which loops back to the model.
The three discovery harnesses: general interaction, state-guided execution, and outcome-guided repair. The same base model runs under all three (RSR paper, Figure 3).

The three are complementary, and the paper measures it well. On a pool of 2,929 terminal tasks (2,509 self-built SWR tasks across about 50 domains plus 420 harder RST tasks), the model ran 14,598 discovery rollouts and passed 2,001 of them, 13.7% (reported). Those passes cover 759 distinct tasks. The best single harness, Terminus 2, covers 565; the union adds 194, a 34.3% relative increase (reported; 194 / 565 = 34.3%, reasoned). Of the 759, 288 are solved by exactly one harness (129 Terminus 2, 91 RSRT, 68 StateM), 163 by two, 308 by all three (reported; the regions in Figure 4b sum to 759, measured).

Left: grouped bar chart of the share of tasks solved by Terminus 2, RSRT, StateM and their union on RST (420), SWR (2,509) and pooled (2,929); the union is highest everywhere, 30.7, 25.1 and 25.9 percent. Right: an upset plot of solve regions on the pooled set: 308 tasks solved by all three, 129 by Terminus 2 only, 91 by RSRT only, 78 by Terminus 2 and RSRT, 68 by StateM only, 50 by Terminus 2 and StateM, 35 by RSRT and StateM.
Coverage of each harness and their union, and the exclusive and overlapping solve regions. Union 759, best single 565, additional 194 (RSR paper, Figure 4).

The raw pool has unequal rollout counts (5,405 Terminus 2, 3,777 RSRT, 5,416 StateM; reported, RSRT had more sandbox errors), so the authors also check a balanced subset of 1,247 tasks with equal attempts per harness. There RSRT is the strongest single harness at 285 tasks, the union reaches 352, and each harness still has exclusive solves: 30, 56 and 21 (reported). The complementarity is not an artefact of sampling.

So a model with three loops can solve a quarter more tasks than a model with one. The catch is what those extra successes look like. A StateM trajectory is full of StateM: phase resets, transition checks, commands like statem goto implement. An RSRT trajectory contains the verifier's "failed, keep going" messages and runs past the point where the model itself would have stopped: RSRT averages 2.20 completion claims per passing trajectory against 1.11 for StateM, and 4.5% of its passes come after a rejected completion (reported). Train on those as-is and you teach the model to expect a controller that will not be there at inference time.

The paper has a direct measurement of what that costs. Direct SFT, fine-tuning on the passing source rollouts with the harness-specific insertion phrases stripped out, lowers Terminal-Bench 2 pass@3 from 57.0% to 53.4% and its per-run mean from 51.7% to 43.8% (reported). Inspecting three trajectories from that checkpoint, the authors saw a recurring dead end: the model loops on the same action pattern without progress. RSRT produces many trajectories longer than 100 steps, and the paper's reading is that a 27B model imitating those cannot sustain them without the harness that kept prodding it (reported; three inspected trajectories is an anecdote, and the paper says as much).

The rewrite: re-solve it, don't edit it

The key design decision is that RSR never edits a trajectory. Stripping control messages from a StateM run would leave a log with holes where the controller used to push. Instead, the same base model plays three roles and produces a new trajectory from a fresh start.

Three-stage pipeline. Discover: tasks fan out to harness 1, 2, 3, producing more difficult solutions. Passing traces go to Rewrite: a planner writes K runbooks per trace with milestones and checks; a critic does rule checks and model checks for solution, verifier and harness leakage, sending leaks back to the planner to resample with the reason; the approved private guide goes to the executor under the general harness, acting in a fresh sandbox, M times per runbook. Rewritten traces go to Verify: task verifier with hidden checks, fail discards, pass goes to a leak audit for underived values, then a filter, producing verified trajectories.
Discover with several harnesses, rewrite with planner, critic and executor (one model in all three roles), then verify and audit before anything becomes training data (RSR paper, Figure 2).

Compact. The source trajectory is cut down to the task instruction, the model's actions and the environment's observations. Harness control messages are removed. This compacted log is input for the planner only; it never becomes training data.

Plan. The planner reads the compacted success and writes a runbook: the required end state, key milestones, useful checks, recovery strategies and common pitfalls. It may describe the task, the interfaces and how to validate. It must not hand over the finished deliverable. It samples K = 4 runbook candidates per source trajectory (reported).

Criticize. Every candidate goes through deterministic checks first (schema validation, removal of known artefacts, rejection of references to tools the general harness doesn't have), then a model-based critic. The critic sees only the public task instruction and the candidate runbook, and decides whether the runbook teaches a procedure or leaks something the executor shouldn't get: the verifier's checks, the solution, harness-specific machinery. A rejected runbook goes back to the planner with the critic's reason, and the planner regenerates. That loop, capped at N iterations in the paper's Figure 1, is the "recursive" in the name.

Execute. For each accepted runbook, the executor solves the task again, in a fresh sandbox, under plain Terminus 2, M = 4 times at temperature 0.7 (reported). The runbook is private guidance: it is in the executor's context while it generates, but it is never written into the trajectory. So the executor has to react to what this sandbox actually prints, rather than replay the old run.

Filter. A rewrite has to pass the task's verifier. Then the model is used once more to flag any value that appears in the trajectory but can't be derived from the task or the environment, a sign that an answer travelled through the runbook, and those rewrites are thrown away. What is kept is the public interaction history: instruction, observations, the model's own responses. The runbook and the critic conversation are dropped.

The widget walks one of the paper's own cases through those stages: an OpenFOAM reconstruction task first solved under StateM in 38 turns. The commands marked "fig 6" are the ones the paper prints for the source and its passing rewrite; the connecting lines are mine, there to show what each stage removes.

One success, rewritten for the general harness
OpenFOAM case · lines marked fig 6 are the paper's, the rest illustrative
discovery harness
StateM trajectory (38 turns, passed)
  • harness[StateM] state=ANALYZE · context reset for phase
  • modelinspect the public scenarios and their summaries
  • harness[StateM] transition check: contract written? yes → IMPLEMENT
  • modelpython3 .../tb_receipt.py validate-contractfig 6
  • modelstatem goto implementfig 6
  • envsummary fields differ across scenarios
  • modelwrite reconstruct.py; rerun feedback
Nothing yet
·

A passing rollout under StateM. Half of what is in it is the harness talking: phase resets, transition checks, and commands that only exist because StateM is running.

Two things make the result trustworthy as training data, at least by construction. Correctness is re-established rather than inherited: every kept rewrite passed the verifier in its own sandbox, so it is a real success under the target harness, not a cleaned-up transcript of someone else's. And the leak defence is layered: the critic screens the runbook before execution, and the underived-value audit screens the trajectory after. Neither is a proof. The critic and the auditor are the same 27B model that is being trained, and the paper reports no false-negative rate for either.

Concept illustration. Left, success experiences: one model under harness A, B and C, each with a passing trace. Middle, recursive planning: the planner writes a runbook with goal, steps, check and avoid lines, and a struck-out solution line flagged as a leak; the critic rejects drafts 1 and 2 back to the planner with the reason, at most N times, and accepts draft 3. Right, re-learn experience: runbooks 1, 2 and 3 each drive several make-test-fix attempts, most passing and one discarded; one success becomes many, up to K times M.
The planner-critic loop is where the recursion lives: a runbook goes back with the critic's reason until it is clean or N attempts run out, and one success becomes up to K × M fresh attempts (RSR paper, Figure 1).

Is it recursive in the self-improvement sense?

No, and the paper is explicit about it. Its related-work section says RSR "does not optimize a single harness or assume an autonomous recursive cycle." There is one round: the base model discovers, the base model rewrites, the base model is fine-tuned once. The improved model is not sent back to rediscover and rewrite again. So the brief question "does the improved model rewrite again?" has a plain answer: not in this paper. A second round is the obvious experiment and nothing here measures it. Compare Dream-RSI, where what recurses across rounds is the search code, and RRSI, where it is the harness; RSR's only loop is inside the data generator.

The training is equally plain: supervised fine-tuning, no RL. The RSR model is trained on the verified rewrites plus the 766 direct passing Terminus 2 trajectories, which need no rewriting because they are already in the target harness (reported). The paper gives no learning rate, epoch count, sequence length or compute budget, so the recipe cannot be reproduced from the text alone.

How many trajectories come out

The rewrite stage targets the StateM and RSRT successes, 636 + 599 = 1,235 source trajectories (reported counts; sum reasoned). With K = 4 runbooks and M = 4 executions each, the ceiling is 16 attempts per source, about 19,760 in total (reasoned). The paper reports 12,893 rewrite rollouts, so roughly a third of runbook slots must have been dropped by the critic or lost to sandbox errors (reasoned). Of those 12,893, 86% pass, giving 11,094 training trajectories (reported; 11,094 / 12,893 = 86.0%, reasoned).

One count does not reconcile. The paper says the rewrites cover 975 unique tasks, but discovery solved only 759 tasks in total, and the StateM and RSRT sources cover at most 630 of them (Table 1's RSRT ∪ StateM row). A rewrite cannot cover a task no source solved. Either 975 counts something else, or the rewrite stage ran on sources beyond those described. I can't resolve it from the paper.

The two case studies give a feel for yield. For a Markdown inline-parsing task from RSRT, 7 of 12 rewrites pass (58.3%), and the median passing rewrite takes 32 turns against the source's 64 (reported). For the OpenFOAM task from StateM, 7 of 8 pass (87.5%), median 29 turns against 38, and the rewrites keep the source's analyse-implement-repair rhythm without any StateM commands (reported). Rewrites guided by the same runbook share more commands than rewrites from different runbooks (19.1 vs 12.1 Jaccard on Markdown, 12.6 vs 10.4 on OpenFOAM; reported), which is the evidence that the runbook steers the procedure without dictating the commands.

Two case studies. Case 1, Markdown inline parsing from an RSRT source, 7 of 12 pass, 64 to 32 median turns; the runbook reads find a compatible parser, match the output summary, probe the table parameter. The source reaches the prebuilt tree-sitter parser at turn 51; the passing rewrite uses it at turn 9; a failing rewrite hits xxd command not found. Case 2, OpenFOAM demo outputs from a StateM source, 7 of 8 pass, 38 to 29 median turns; the source runs tb_receipt.py validate-contract and statem goto implement; the passing rewrite inspects reconstruct.py and rebuilds; a failing rewrite hits a NameError and a KeyError.
A source run, a passing rewrite and a failing rewrite guided by the same runbook, for two tasks. The rewrite under Terminus 2 has no StateM commands in it (RSR paper, Figure 6).

The model card shows three more tasks the same way, with wall-clock time. All three sources ran under continue-until-timeout and hit the limit, and the passing rewrites are less than half as long: 64 → 34 steps and 2.0 h → 22 min on a tree-sitter task, 59 → 27 steps on CMake, 56 → 23 on OpenFOAM (reported). The card calls it a 2-hour limit; the paper says RSRT allows up to 180 minutes. These are hand-picked successes, each with 4 of 4 replays passing; the paper's own 7-of-12 case is the more representative yield.

Horizontal bars comparing an expert trajectory and a passing rewrite for three SWR tasks, coloured by command kind (explore, edit, run, install, wait, other). tree_sitter_markdown_inline_05: expert 64 steps, 153 commands, 2.0 h; rewrite 34 steps, 78 commands, 22 min. cmake_generated_header_config_06: expert 59 steps, 94 commands, 2.0 h; rewrite 27 steps, 58 commands, 21 min. openfoam_scalar_transport_boundary_probe_02: expert 56 steps, 72 commands, 1.9 h; rewrite 23 steps, 51 commands, 28 min.
Expert (base model under continue-until-timeout) versus a passing rewrite (base model under plain Terminus 2, runbook kept private) on three sample SWR tasks; selected examples, not averages (IntelligenceLab/RSR-27B model card).

Results

All three models are evaluated under plain Terminus 2, three runs per benchmark. Pass@3 counts a task if any of the three runs passes; the mean averages the three per-run pass rates.

Benchmark (tasks)MetricBaseDirect SFTRSR
Terminal-Bench 2 (89)pass@357.0%53.4%74.2%
Terminal-Bench 3 (74)pass@30.0%5.4%9.5%
Terminal-Bench 4 (66)pass@31.5%4.5%9.1%
TBH, authors' (100)pass@339.0%56.0%63.0%
SWR100, authors' (100)pass@33.0%3.0%6.0%
Terminal-Bench 2mean of 351.7%43.8%70.1%
TBHmean of 331.0%42.5%54.0%
LHTB, authors' (46)process reward0.210.250.29

All reported (paper, Table 5). RSR beats Direct SFT on every row: by 20.8, 4.1, 4.6, 7.0 and 3.0 points of pass@3 and by 26.3, 5.4, 2.9, 11.5 and 2.0 points of per-run mean on the five pass-rate benchmarks (reported, and the differences check out, reasoned). On Long-Horizon Terminal-Bench, all three complete 0 of 46 tasks; the process-reward gain is partial progress, not solved tasks.

Qwen-3.8-27B on terminal benchmarks
reported · paper Tables 5 and 6
BaseDirect SFTRSR
TB 2 (89 tasks)
57.0%
53.4%
74.2%
TB 3 (74 tasks)
0.0%
5.4%
9.5%
TB 4 (66 tasks)
1.5%
4.5%
9.1%
TBH (100 tasks)
39.0%
56.0%
63.0%
SWR100 (100 tasks)
3.0%
3.0%
6.0%

All three models evaluated under plain Terminus 2. Direct SFT on the raw multi-harness successes loses 3.6 points on TB 2; RSR gains 17.2 over Base.

What the numbers do and don't show

The headline holds on the public benchmark. Terminal-Bench 2 is not the authors' own, and +17.2 points of pass@3 and +18.4 of per-run mean over the base model (reasoned from Table 5) is a large move for one round of SFT on self-generated data. The contrast with Direct SFT, which went down, is the paper's best evidence that the rewrite step matters and not just the extra successes.

It is not a controlled data comparison. Direct SFT trains on the raw passing rollouts; the paper doesn't give its count, but the pool it describes holds 2,001. RSR trains on 11,094 rewrites plus 766 direct trajectories, 11,860 in all (reported counts, sum reasoned), about 5.9 times as many (reasoned). The paper has no size-matched ablation, and no "RSR without the critic" or "rewrite only Terminus 2" arm. So I can't separate "cleaner trajectories" from "six times more trajectories" in the 20.8-point gap. Some of both, presumably.

Two of the benchmarks are the authors' own, and one is close to the training data. TBH comes from the same RST pipeline (Li et al., 2026a) that supplied 420 of the training tasks. The paper doesn't describe any overlap check between the TBH tasks and the training pool. The TBH gains may be fine, but they carry less weight than the Terminal-Bench 2 gains.

The pass@3 counts are not exact fractions. Table 5 prints "57.0% (51/89)" and notes that the counts are back-computed from the percentages. But 51/89 is 57.3%, and no integer over 89 gives 57.0% or 53.4% (reasoned). The percentages come from somewhere other than a plain count over 89 tasks, perhaps tasks that failed to run, and the paper doesn't say. It doesn't change the story; it is a reason to read the decimals loosely.

The harness alone gets you most of the way, on their own benchmark. Table 6 runs the untrained base model once under each discovery harness. On Terminal-Bench 2 that gives 51.7% (Terminus 2), 55.1% (RSRT), 58.4% (StateM). On TBH it gives 33.0%, 66.0% and 46.0% (reported). The untrained model under RSRT, at 66.0% in a single run, beats the trained RSR model under Terminus 2 on TBH: 63.0% pass@3, 54.0% per run. On Terminal-Bench 2 the trained model wins clearly, 70.1% per run against StateM's 58.4%. That is the trade the paper is making explicit: RSR is worth it when you want one cheap, general loop at deployment, not the best possible score on a task family where a heavier loop already works. (Table 6 also prints the TB3 union as "4.3 (3)"; 3/74 is 4.1%, reasoned.)

The released model

IntelligenceLab/RSR-27B is a Qwen3_5ForConditionalGeneration checkpoint, 64 layers, hidden size 5,120, with 55.56 GB of safetensors across 16 main shards plus two named model-missing-from-origin-* (measured from config.json, the file listing and model.safetensors.index.json). Those two hold the vision tower (333 tensors) and the multi-token prediction head (15 tensors), measured. My reading is that the fine-tune touched the language model and the rest was copied back in from the base checkpoint; the card doesn't say. 55.56 GB at 2 bytes per parameter is about 27.8B parameters (reasoned, assuming BF16). The card has no evaluation table, only the three sample trajectories above and a link to a trajectory viewer.

What I take from it

The mechanism is simple enough to copy, and the insight underneath it is the useful part: a success under a heavy harness is a hint, not a demonstration. You don't train on it; you have the model solve the task again, under the loop it will actually run in, with the hint kept out of the log. Correctness comes back from the verifier, not from the original run.

What it leaves open is everything a second round would answer. Does RSR-27B under the three harnesses discover tasks the base model couldn't, and does rewriting those help again, or does the gain flatten after one pass? The name promises recursion at that level, and this paper only does it inside the runbook loop. It also leaves the boundary where a heavier loop still beats a trained model, which Table 6 suggests is real, unmapped.

Related on the site: Scaling agentic RL and MiMo-V2.6's RL environments on where verifiable terminal tasks come from, and Qwen3.8-Max for the family RSR's base model belongs to.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Recursive Self-Rewrite: teach the model what the harness did", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026recursiveselfrewrite,
  author = {Satyajit Ghana},
  title  = {Recursive Self-Rewrite: teach the model what the harness did},
  url    = {https://ai.thesatyajit.com/articles/recursive-self-rewrite},
  year   = {2026}
}
share