2026-10-06 · 17 min · explainer · agents · llm · fine-tuning · self-improvement · benchmarks · qwen · open-weights
The site keeps returning to one uncomfortable fact about agents: the loop around the model decides a lot of the score. Agent harnesses made the case in general, the harness effect put a token bill on it, and the harness is the generalizer argued the loop can carry generalization the weights don't. Hugging Face's multi-harness RL guide makes the same point from the training side: the same weights score 62.1% under Mini-SWE-Agent and 33.2% under Claude Code.
Most of the work that follows from that fact improves the harness and leaves the model alone. RRSI and Recursive Harness Self-Improvement evolve the loop; Dream-RSI evolves a scheduler. Recursive Self-Rewrite (RSR; Li et al., arXiv 2610.02826, 2 Oct 2026) goes the other way. It uses several harnesses as discovery tools, collects whatever they help the model solve, and then moves that skill into the weights, so the model can do it again under a plain loop with none of that help. The released checkpoint is IntelligenceLab/RSR-27B on Hugging Face, Apache-2.0.
Every number below is labelled. Reported is the paper's or the model card's figure, not re-run. Measured is something I read off a file myself. Reasoned is my arithmetic on the other two. I did not train or evaluate anything.
The problem: a success belongs to the model and the loop together
A harness decides how the model sees observations, calls tools, tracks progress, checks its work, recovers and stops. Change it and you change which tasks get solved. RSR uses three, all driving the same Qwen-3.8-27B on terminal tasks:
- Terminus 2, the general loop from the Harbor framework. The model issues shell commands, reads the output, and decides when it is done. This is also the target harness: the one the model will be deployed and evaluated under.
- StateM, which organizes a run around persistent workflow states, phase-local context and checked transitions: write a contract, pass a check, move to the next phase.
- RSRT (Recursive Self-Reflect Terminus), the authors' modification of Terminus 2. When the model claims it is done, the harness runs the task verifier. On a fail it reports only "failed", never the verifier's internals, and tells the model to review its trajectory, debug and continue, up to 180 minutes of wall clock per task (reported).

The three are complementary, and the paper measures it well. On a pool of 2,929 terminal tasks (2,509 self-built SWR tasks across about 50 domains plus 420 harder RST tasks), the model ran 14,598 discovery rollouts and passed 2,001 of them, 13.7% (reported). Those passes cover 759 distinct tasks. The best single harness, Terminus 2, covers 565; the union adds 194, a 34.3% relative increase (reported; 194 / 565 = 34.3%, reasoned). Of the 759, 288 are solved by exactly one harness (129 Terminus 2, 91 RSRT, 68 StateM), 163 by two, 308 by all three (reported; the regions in Figure 4b sum to 759, measured).

The raw pool has unequal rollout counts (5,405 Terminus 2, 3,777 RSRT, 5,416 StateM; reported, RSRT had more sandbox errors), so the authors also check a balanced subset of 1,247 tasks with equal attempts per harness. There RSRT is the strongest single harness at 285 tasks, the union reaches 352, and each harness still has exclusive solves: 30, 56 and 21 (reported). The complementarity is not an artefact of sampling.
So a model with three loops can solve a quarter more tasks than a model with one. The catch is what
those extra successes look like. A StateM trajectory is full of StateM: phase resets, transition
checks, commands like statem goto implement. An RSRT trajectory contains the verifier's
"failed, keep going" messages and runs past the point where the model itself would have stopped:
RSRT averages 2.20 completion claims per passing trajectory against 1.11 for StateM, and 4.5% of
its passes come after a rejected completion (reported). Train on those as-is and you teach the
model to expect a controller that will not be there at inference time.
The paper has a direct measurement of what that costs. Direct SFT, fine-tuning on the passing source rollouts with the harness-specific insertion phrases stripped out, lowers Terminal-Bench 2 pass@3 from 57.0% to 53.4% and its per-run mean from 51.7% to 43.8% (reported). Inspecting three trajectories from that checkpoint, the authors saw a recurring dead end: the model loops on the same action pattern without progress. RSRT produces many trajectories longer than 100 steps, and the paper's reading is that a 27B model imitating those cannot sustain them without the harness that kept prodding it (reported; three inspected trajectories is an anecdote, and the paper says as much).
The rewrite: re-solve it, don't edit it
The key design decision is that RSR never edits a trajectory. Stripping control messages from a StateM run would leave a log with holes where the controller used to push. Instead, the same base model plays three roles and produces a new trajectory from a fresh start.

Compact. The source trajectory is cut down to the task instruction, the model's actions and the environment's observations. Harness control messages are removed. This compacted log is input for the planner only; it never becomes training data.
Plan. The planner reads the compacted success and writes a runbook: the required end state, key milestones, useful checks, recovery strategies and common pitfalls. It may describe the task, the interfaces and how to validate. It must not hand over the finished deliverable. It samples K = 4 runbook candidates per source trajectory (reported).
Criticize. Every candidate goes through deterministic checks first (schema validation, removal of known artefacts, rejection of references to tools the general harness doesn't have), then a model-based critic. The critic sees only the public task instruction and the candidate runbook, and decides whether the runbook teaches a procedure or leaks something the executor shouldn't get: the verifier's checks, the solution, harness-specific machinery. A rejected runbook goes back to the planner with the critic's reason, and the planner regenerates. That loop, capped at N iterations in the paper's Figure 1, is the "recursive" in the name.
Execute. For each accepted runbook, the executor solves the task again, in a fresh sandbox, under plain Terminus 2, M = 4 times at temperature 0.7 (reported). The runbook is private guidance: it is in the executor's context while it generates, but it is never written into the trajectory. So the executor has to react to what this sandbox actually prints, rather than replay the old run.
Filter. A rewrite has to pass the task's verifier. Then the model is used once more to flag any value that appears in the trajectory but can't be derived from the task or the environment, a sign that an answer travelled through the runbook, and those rewrites are thrown away. What is kept is the public interaction history: instruction, observations, the model's own responses. The runbook and the critic conversation are dropped.
The widget walks one of the paper's own cases through those stages: an OpenFOAM reconstruction task first solved under StateM in 38 turns. The commands marked "fig 6" are the ones the paper prints for the source and its passing rewrite; the connecting lines are mine, there to show what each stage removes.
- harness[StateM] state=ANALYZE · context reset for phase
- modelinspect the public scenarios and their summaries
- harness[StateM] transition check: contract written? yes → IMPLEMENT
- modelpython3 .../tb_receipt.py validate-contractfig 6
- modelstatem goto implementfig 6
- envsummary fields differ across scenarios
- modelwrite reconstruct.py; rerun feedback
A passing rollout under StateM. Half of what is in it is the harness talking: phase resets, transition checks, and commands that only exist because StateM is running.
Two things make the result trustworthy as training data, at least by construction. Correctness is re-established rather than inherited: every kept rewrite passed the verifier in its own sandbox, so it is a real success under the target harness, not a cleaned-up transcript of someone else's. And the leak defence is layered: the critic screens the runbook before execution, and the underived-value audit screens the trajectory after. Neither is a proof. The critic and the auditor are the same 27B model that is being trained, and the paper reports no false-negative rate for either.

Is it recursive in the self-improvement sense?
No, and the paper is explicit about it. Its related-work section says RSR "does not optimize a single harness or assume an autonomous recursive cycle." There is one round: the base model discovers, the base model rewrites, the base model is fine-tuned once. The improved model is not sent back to rediscover and rewrite again. So the brief question "does the improved model rewrite again?" has a plain answer: not in this paper. A second round is the obvious experiment and nothing here measures it. Compare Dream-RSI, where what recurses across rounds is the search code, and RRSI, where it is the harness; RSR's only loop is inside the data generator.
The training is equally plain: supervised fine-tuning, no RL. The RSR model is trained on the verified rewrites plus the 766 direct passing Terminus 2 trajectories, which need no rewriting because they are already in the target harness (reported). The paper gives no learning rate, epoch count, sequence length or compute budget, so the recipe cannot be reproduced from the text alone.
How many trajectories come out
The rewrite stage targets the StateM and RSRT successes, 636 + 599 = 1,235 source trajectories (reported counts; sum reasoned). With K = 4 runbooks and M = 4 executions each, the ceiling is 16 attempts per source, about 19,760 in total (reasoned). The paper reports 12,893 rewrite rollouts, so roughly a third of runbook slots must have been dropped by the critic or lost to sandbox errors (reasoned). Of those 12,893, 86% pass, giving 11,094 training trajectories (reported; 11,094 / 12,893 = 86.0%, reasoned).
One count does not reconcile. The paper says the rewrites cover 975 unique tasks, but discovery solved only 759 tasks in total, and the StateM and RSRT sources cover at most 630 of them (Table 1's RSRT ∪ StateM row). A rewrite cannot cover a task no source solved. Either 975 counts something else, or the rewrite stage ran on sources beyond those described. I can't resolve it from the paper.
The two case studies give a feel for yield. For a Markdown inline-parsing task from RSRT, 7 of 12 rewrites pass (58.3%), and the median passing rewrite takes 32 turns against the source's 64 (reported). For the OpenFOAM task from StateM, 7 of 8 pass (87.5%), median 29 turns against 38, and the rewrites keep the source's analyse-implement-repair rhythm without any StateM commands (reported). Rewrites guided by the same runbook share more commands than rewrites from different runbooks (19.1 vs 12.1 Jaccard on Markdown, 12.6 vs 10.4 on OpenFOAM; reported), which is the evidence that the runbook steers the procedure without dictating the commands.

The model card shows three more tasks the same way, with wall-clock time. All three sources ran under continue-until-timeout and hit the limit, and the passing rewrites are less than half as long: 64 → 34 steps and 2.0 h → 22 min on a tree-sitter task, 59 → 27 steps on CMake, 56 → 23 on OpenFOAM (reported). The card calls it a 2-hour limit; the paper says RSRT allows up to 180 minutes. These are hand-picked successes, each with 4 of 4 replays passing; the paper's own 7-of-12 case is the more representative yield.

Results
All three models are evaluated under plain Terminus 2, three runs per benchmark. Pass@3 counts a task if any of the three runs passes; the mean averages the three per-run pass rates.
| Benchmark (tasks) | Metric | Base | Direct SFT | RSR |
|---|---|---|---|---|
| Terminal-Bench 2 (89) | pass@3 | 57.0% | 53.4% | 74.2% |
| Terminal-Bench 3 (74) | pass@3 | 0.0% | 5.4% | 9.5% |
| Terminal-Bench 4 (66) | pass@3 | 1.5% | 4.5% | 9.1% |
| TBH, authors' (100) | pass@3 | 39.0% | 56.0% | 63.0% |
| SWR100, authors' (100) | pass@3 | 3.0% | 3.0% | 6.0% |
| Terminal-Bench 2 | mean of 3 | 51.7% | 43.8% | 70.1% |
| TBH | mean of 3 | 31.0% | 42.5% | 54.0% |
| LHTB, authors' (46) | process reward | 0.21 | 0.25 | 0.29 |
All reported (paper, Table 5). RSR beats Direct SFT on every row: by 20.8, 4.1, 4.6, 7.0 and 3.0 points of pass@3 and by 26.3, 5.4, 2.9, 11.5 and 2.0 points of per-run mean on the five pass-rate benchmarks (reported, and the differences check out, reasoned). On Long-Horizon Terminal-Bench, all three complete 0 of 46 tasks; the process-reward gain is partial progress, not solved tasks.
All three models evaluated under plain Terminus 2. Direct SFT on the raw multi-harness successes loses 3.6 points on TB 2; RSR gains 17.2 over Base.
What the numbers do and don't show
The headline holds on the public benchmark. Terminal-Bench 2 is not the authors' own, and +17.2 points of pass@3 and +18.4 of per-run mean over the base model (reasoned from Table 5) is a large move for one round of SFT on self-generated data. The contrast with Direct SFT, which went down, is the paper's best evidence that the rewrite step matters and not just the extra successes.
It is not a controlled data comparison. Direct SFT trains on the raw passing rollouts; the paper doesn't give its count, but the pool it describes holds 2,001. RSR trains on 11,094 rewrites plus 766 direct trajectories, 11,860 in all (reported counts, sum reasoned), about 5.9 times as many (reasoned). The paper has no size-matched ablation, and no "RSR without the critic" or "rewrite only Terminus 2" arm. So I can't separate "cleaner trajectories" from "six times more trajectories" in the 20.8-point gap. Some of both, presumably.
Two of the benchmarks are the authors' own, and one is close to the training data. TBH comes from the same RST pipeline (Li et al., 2026a) that supplied 420 of the training tasks. The paper doesn't describe any overlap check between the TBH tasks and the training pool. The TBH gains may be fine, but they carry less weight than the Terminal-Bench 2 gains.
The pass@3 counts are not exact fractions. Table 5 prints "57.0% (51/89)" and notes that the counts are back-computed from the percentages. But 51/89 is 57.3%, and no integer over 89 gives 57.0% or 53.4% (reasoned). The percentages come from somewhere other than a plain count over 89 tasks, perhaps tasks that failed to run, and the paper doesn't say. It doesn't change the story; it is a reason to read the decimals loosely.
The harness alone gets you most of the way, on their own benchmark. Table 6 runs the untrained base model once under each discovery harness. On Terminal-Bench 2 that gives 51.7% (Terminus 2), 55.1% (RSRT), 58.4% (StateM). On TBH it gives 33.0%, 66.0% and 46.0% (reported). The untrained model under RSRT, at 66.0% in a single run, beats the trained RSR model under Terminus 2 on TBH: 63.0% pass@3, 54.0% per run. On Terminal-Bench 2 the trained model wins clearly, 70.1% per run against StateM's 58.4%. That is the trade the paper is making explicit: RSR is worth it when you want one cheap, general loop at deployment, not the best possible score on a task family where a heavier loop already works. (Table 6 also prints the TB3 union as "4.3 (3)"; 3/74 is 4.1%, reasoned.)
The released model
IntelligenceLab/RSR-27B is a Qwen3_5ForConditionalGeneration
checkpoint, 64 layers, hidden size 5,120, with 55.56 GB of safetensors across 16 main shards plus two
named model-missing-from-origin-* (measured from config.json, the file listing and
model.safetensors.index.json). Those two hold the vision tower (333 tensors) and the multi-token
prediction head (15 tensors), measured. My reading is that the fine-tune touched the language model and
the rest was copied back in from the base checkpoint; the card doesn't say. 55.56 GB at 2 bytes per
parameter is about 27.8B parameters (reasoned, assuming BF16). The card has no evaluation table, only
the three sample trajectories above and a link to a trajectory viewer.
What I take from it
The mechanism is simple enough to copy, and the insight underneath it is the useful part: a success under a heavy harness is a hint, not a demonstration. You don't train on it; you have the model solve the task again, under the loop it will actually run in, with the hint kept out of the log. Correctness comes back from the verifier, not from the original run.
What it leaves open is everything a second round would answer. Does RSR-27B under the three harnesses discover tasks the base model couldn't, and does rewriting those help again, or does the gain flatten after one pass? The name promises recursion at that level, and this paper only does it inside the runbook loop. It also leaves the boundary where a heavier loop still beats a trained model, which Table 6 suggests is real, unmapped.
Related on the site: Scaling agentic RL and MiMo-V2.6's RL environments on where verifiable terminal tasks come from, and Qwen3.8-Max for the family RSR's base model belongs to.