# RecursiveMAS: a multi-agent system folded into one looped transformer

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/recursivemas
> date: 2026-10-02
> tags: llm, agents, multi-agent, looped-transformers, recurrent-depth, latent-reasoning, inference-optimization

A [looped transformer](/articles/looped-transformers-matched-compute) buys reasoning depth by reusing the
same layers over and over instead of stacking new ones. You run the block, feed its output back as input,
run it again — recurrent depth in latent space, no new parameters. [LOTUS](/articles/lotus-latent-reasoning)
does the single-model version of this for chain-of-thought: it reasons in hidden states across $R$ looped
passes rather than writing every step out as tokens.

RecursiveMAS asks the obvious next question. A multi-agent system (MAS) is already a stack of LLM calls — a
Planner, then a Solver, then a Critic — wired in sequence. What if you treat **that whole pipeline as one
looped transformer**, where each agent is a layer, the states passed between agents are latent vectors
instead of text, and the last agent's output loops back into the first for another round? Only the final
round decodes to words. Everything in between stays in embedding space.

That is the entire idea, and it is a genuinely different lever from the two that dominate agent research:
not a bigger model, not more agents, but **deeper connected collaboration** over the same team. The paper
(["Recursive Multi-Agent Systems"](https://arxiv.org/abs/2604.25917), Zou et al., NeurIPS 2026) builds it
on small open models, freezes all of them, and trains a 13-million-parameter connector that is 0.31% of the
weights. The headline claims are +8.3% average accuracy over the strongest baseline, a 1.2×–2.4×
end-to-end speedup, and a 34.6%–75.6% token reduction. Below I walk the mechanism from first principles,
then check each of those numbers against the paper's tables — including one metric choice that inflates the
most eye-catching rows.

<RepoCard repo="RecursiveMAS/RecursiveMAS" />

## The usual way: agents talk in text

A standard MAS is a group of LLMs that pass natural language to each other. The Planner writes a plan in
words; the Solver reads those words, re-tokenizes them, and writes a solution; the Critic reads *that* and
writes feedback. Every hand-off is a full round trip through the vocabulary: decode hidden states to
discrete tokens, then embed those tokens back into hidden states at the next agent.

Two costs come out of this. The obvious one is **tokens** — every intermediate message is generated one
token at a time, which is where [inference latency lives](/articles/how-llm-inference-works), and a
recursive MAS that loops the whole team re-pays that cost every round. The subtler one is **information
loss**: decoding to a single token sequence collapses a rich distribution over the vocabulary into one
sampled path, and the next agent only ever sees the path, not the distribution behind it.

RecursiveMAS keeps the collaboration structure — the same Planner/Solver/Critic topology — but replaces the
text channel between agents with a latent one. Scrub the recursion depth and flip the channel to see the
two effects that produces:

<RecursionStepper />

## Agents as layers: the inner and outer links

To make "each agent is a layer" literal, you need a way to turn one agent's hidden states into the next
agent's inputs without going through text. That connector is **RecursiveLink**, and it comes in two flavours
for the two kinds of hand-off in the system.

The first hand-off is *within* an agent. An ordinary LLM generates autoregressively: produce a token, embed
it, feed it back in for the next step. RecursiveMAS wants the agent to keep thinking without emitting a
token, so it takes the agent's **last-layer** embedding $h$ and maps it back into the **input-layer**
embedding space for the next forward step. The paper calls the result an *ongoing latent thought*. That map
is the **inner link** $\mathcal{R}_{\text{in}}$:

$$
\mathcal{R}_{\text{in}}(h) = h + W_2\,\sigma(W_1 h)
$$

where $W_1, W_2$ are linear layers and $\sigma$ is GELU. It is a two-layer MLP with a residual connection —
tiny. The agent runs this for $m$ steps, producing a short continuous sequence of latent thoughts
$[h_t, h_{t+1}, \dots, h_{t+m}]$ instead of $m$ decoded tokens.

The second hand-off is *across* agents, and heterogeneous agents do not share a hidden width — a 1.7B
Planner and a 1.5B Solver have different $d_h$. The **outer link** $\mathcal{R}_{\text{out}}$ handles that
with one extra linear layer $W_3$ on the residual branch to project the source width to the target width:

$$
\mathcal{R}_{\text{out}}(h) = W_3 h + W_2\,\sigma(W_1 h)
$$

<Figure
  src="https://ai.thesatyajit.com/articles/recursivemas/fig2.png"
  alt="Two side-by-side diagrams of the RecursiveLink module. Left (Inner): a Last-layer Emb feeds a stack of Linear, GELU, Linear, whose output is summed at a plus node with a residual connection carrying the Last-layer Emb directly, producing an Input-layer Emb for the same agent. Right (Outer): an Agent A_i Emb feeds the same Linear, GELU, Linear stack, but the residual branch now passes through an additional Linear block before the plus node, producing an Agent A_j Emb in the target agent's space."
  caption="RecursiveLink is a two-layer residual MLP. The inner link maps an agent's last-layer embedding back to its own input space; the outer link adds a linear W3 on the residual branch to map between agents of different hidden widths (RecursiveMAS, Figure 3)."
/>

Why the residual term in both? Because the residual carries the original embedding through largely
unchanged, so the MLP only has to learn the *difference* between the two embedding spaces, not the whole
projection from scratch. That is the same reason residual connections stabilize deep nets, applied to a
cross-model adapter: the link aligns distributions rather than relearning semantics, which the authors argue
makes training more stable. The branches are small enough that the two links across the whole system come to
just **13.12M trainable parameters** (reported, Table 5).

## Closing the loop

Chain the agents and you get the full architecture. Agent $A_1$ generates its latent thoughts through its
inner link, the outer link projects them into $A_2$'s input space, $A_2$ conditions on both its own prompt
and $A_1$'s transferred latents, and so on down the line. When the last agent $A_N$ finishes, its latent
output — the system's current latent answer — is passed **back to $A_1$** through the inner-outer link,
closing a recurrent loop over the whole team. Each new round lets every agent reflect on the previous
round's system output and refine. Only the last agent, in the final round $n$, actually decodes to text.

<Figure
  src="https://ai.thesatyajit.com/articles/recursivemas/fig1.png"
  alt="The RecursiveMAS architecture as a left-to-right pipeline. Agent A1 produces a row of latent thoughts (last-layer embeddings h_t, h_{t+1}, ...) that are routed by an inner link back into its own input embeddings. Those latents are then passed through an outer link to Agent A2, whose inputs are the A2-aligned embeddings conditioned on A1's output; A2 runs its own inner link. This continues to Agent A_N, which decodes the final outputs in the last recursion round. A looping symbol marks that the last agent feeds back to the first for rounds before the last."
  caption="Each agent generates latent thoughts with its inner link, then transfers them to the next agent through the outer link. After the last agent, its latents feed back to the first, forming a recursive loop; only the final round decodes to text (RecursiveMAS, Figure 2)."
/>

The mapping to a looped transformer is now exact: agents are layers, a recursion round is one pass through
the stack, and looping the last layer back to the first is the recurrence. The difference from a
single-model [looped transformer](/architectures/looped-transformer) is only that the "layers" are whole
frozen LLMs of different sizes, glued by trained links.

## Why latent transfer is actually cheaper

It would be easy to assume the latent channel is faster just because "latents are smaller than text." The
paper makes the sharper argument through runtime complexity. For a text-based recursive MAS with the same
loop structure, each round costs (reported):

$$
\Theta\!\big(N(\,m|V|d_h + (t{+}m)d_h^2 + (t{+}m)^2 d_h\,)\big)
$$

The term that matters is $m|V|d_h$: generating $m$ intermediate tokens means projecting hidden states onto
the full vocabulary $|V|$ (tens of thousands wide) at every step, for every agent, every round. RecursiveMAS
removes that projection — it never decodes intermediate steps — and the term collapses to $m d_h^2$:

$$
\Theta\!\big(N(\,m d_h^2 + (t{+}m)d_h^2 + (t{+}m)^2 d_h\,)\big)
$$

Since $|V| \gg d_h$ for these small models (a 150k-token vocabulary against a hidden size of a couple
thousand), trading $m|V|d_h$ for $m d_h^2$ is a real saving, and it compounds: it is paid $N$ times per round
and once per recursion round. That is the whole reason the token count falls and the wall-clock speedup
*grows* with recursion depth — deeper loops would have re-decoded more intermediate text, so skipping it
saves more. The efficiency is structural, not a tuning artifact.

<Callout type="note">
The catch is that the latent channel is not human-readable. In a text MAS you can log what the Planner told
the Solver; here the hand-off is a vector. The paper does not offer a way to inspect or audit what agents
pass to each other, and several of the first questions under the authors' announcement were exactly this —
how do you monitor a system whose agents communicate in latent states? Treat interpretability of the inner
channel as an open cost of the method, not a solved part of it.
</Callout>

## Training: freeze the models, train the wiring

Because the base LLMs are frozen, there is very little to optimize — only the links — but *how* you optimize
them across a loop matters. The recipe is a two-stage **inner-outer loop** (reported).

The **inner loop** warms up each agent on its own, in parallel. For a training pair $(x, y)$, it passes the
ground-truth text $y$ through the agent's *input embedding layer* to get a target, then trains that agent's
inner link so its generated latent thoughts align with that target under a cosine-similarity regression. In
effect, each agent learns to "think" in latent form what it would otherwise have written in text, without
the decode-and-re-encode step.

The **outer loop** then trains the system as one entity. It unrolls the full looped MAS for $n$ rounds and
backpropagates through the entire recursive trace, so gradients flow across rounds and the outer links get
*shared credit* — every link is updated for its contribution to the final answer, several rounds downstream.
This is the part that makes it "whole-system co-optimization" rather than training agents in isolation.

There is a non-obvious reason the latent loop trains more stably than you might fear. The authors show that
if you tried to do this through text with standard supervised fine-tuning, confident tokens (entropy
$\le \epsilon$ with $\epsilon \ll 1$ — i.e. a near-one-hot softmax) drive the gradient through the decoding
step toward zero: **gradient vanishing** across the loop. RecursiveLink skips the softmax, and the paper
proves its gradient norm stays close to 1 through the looped backpropagation (reported; a theoretical result
under stated assumptions). Looping through discrete text is where recursive training breaks; looping through
latents is where it keeps working.

## Checking the headline numbers

The paper instantiates RecursiveMAS under four collaboration patterns — sequential, mixture, distillation,
and deliberation — across nine benchmarks in math (MATH500, AIME2025, AIME2026), science and medicine
(GPQA-Diamond, MedQA), code (LiveCodeBench-v6, MBPP+), and search QA (HotpotQA, Bamboogle). The sequential
team I drew above is real and on Hugging Face: a Qwen3-1.7B Planner, a Qwen2.5-Math-1.5B Solver, and a
Llama3.2-1B Critic for the "light" setting, scaled up to Gemma3-4B / Qwen3.5-4B / Llama3.2-3B.

**Does accuracy scale with recursion depth?** Yes, and this is the cleanest result. Averaged over the
math/science/code tasks, RecursiveMAS beats the text-based recursive baseline by +3.4 points at $r=1$, +6.0
at $r=2$, and +7.2 at $r=3$ (reported, Table 2), with the relative margin widening from 8.1% to 20.2% as the
loop deepens. Efficiency scales the same direction: the end-to-end speedup grows 1.2× → 1.9× → 2.4× and the
token reduction grows 34.6% → 65.5% → 75.6% across $r = 1, 2, 3$ (reported). More recursion is more
advantage on both axes at once — the point of the stepper above.

The scaling has two knobs, not one: recursion depth at *training* time and at *inference* time are
separate, and the paper varies both. Accuracy climbs toward the corner where both are large — training
teaches the system to form refinement-ready latent states, and inference recursion then cashes that
structure in for test-time gains. The same figure shows the framework is not tied to the sequential
topology: it lifts the best standalone specialist under mixture, deliberation, and distillation patterns too.

<Figure
  src="https://ai.thesatyajit.com/articles/recursivemas/fig3.png"
  alt="Top row: four heatmaps titled RecursiveMAS Scaling Law for MATH500, AIME2025, GPQA-D, and Code Gen, each a grid of training recursion round (x-axis, 1 to 4) against inference recursion round (y-axis, 1 to 4). Only the lower-right triangle is filled, and accuracy increases toward the top-right corner where both training and inference recursion are deep. Bottom row: three grouped bar charts titled Collaboration Patterns — Mixture-Style, Deliberation-Style, and Distillation-Style — where the dark RecursiveMAS bar is highest in almost every benchmark group against the standalone specialist agents, with per-benchmark end-to-end speedups around 1.4 to 1.6 times annotated on the distillation chart."
  caption="RecursiveMAS scales along two axes at once — training-time and inference-time recursion depth — with accuracy highest where both are deep (top), and the gains hold across mixture, deliberation, and distillation collaboration patterns, not just the sequential one (bottom) (RecursiveMAS, Figure 1)."
/>

**The +8.3% whole-system claim.** At $r=3$, against a broad set of baselines (Table 3):

| Method | MATH500 | AIME2025 | AIME2026 | GPQA-D | LiveCodeBench | MedQA |
|---|---|---|---|---|---|---|
| Single Agent (Full-SFT) | 83.2 | 73.3 | 76.7 | 62.8 | 38.6 | 77.0 |
| Mixture-of-Agents | 79.8 | 60.0 | 63.3 | 47.6 | 27.0 | 57.5 |
| TextGrad | 84.9 | 73.3 | 76.7 | 62.5 | 39.8 | 77.2 |
| LoopLM | 84.6 | 66.7 | 63.3 | 48.1 | 24.9 | 56.4 |
| Recursive-TextMAS | 85.8 | 73.3 | 73.3 | 61.6 | 38.7 | 77.0 |
| **RecursiveMAS** | **88.0** | **86.7** | **86.7** | **66.2** | **42.9** | **79.3** |

RecursiveMAS wins every column — that part holds. But the +8.3% figure deserves a second look. The paper
describes it as the average improvement "over the strongest baseline on each benchmark" across all nine
benchmarks. On the six benchmarks shown here, if I take each column's strongest baseline and average the
gap, I get **+5.7 points** (reasoned, from Table 3); RecursiveMAS's six-benchmark mean of ~75.0 sits about
**+5.9 points** over the best overall baseline, TextGrad at ~69.1 (reasoned). The headline +8.3% is a
nine-benchmark number, so the two search-QA sets and MBPP+ — not in this table — must carry larger margins
to lift the average. The claim is the paper's (reported); I can confirm the direction and the per-column
wins, but the +8.3 specifically is not reproducible from the comparison table alone.

<Callout type="warning">
The AIME rows are **Pass@10**, not Pass@1 — the paper reports AIME2025/2026 as Pass@10 "for testing
robustness." An 86.7 on AIME under Pass@10 means the correct answer appears in at least one of ten samples,
which is a very different quantity from solving it in one shot. The biggest-looking gains in the table
(AIME jumping from 73.3 to 86.7) are on exactly these Pass@10 rows, so read them as the paper labels them,
not as single-shot accuracy.
</Callout>

**The cost table.** This is the efficiency argument that most interests me, because it is measured on the
same scaled sequential setup, not estimated (reported, Table 5):

| Method | Peak GPU mem | Trainable params | Est. cost | Avg. acc |
|---|---|---|---|---|
| LoRA | 21.67 GB | 15.92M (0.37%) | \$6.64 | 66.9 |
| Full-SFT | 41.40 GB | 4.21B (100%) | \$9.67 | 68.6 |
| **RecursiveMAS** | **15.29 GB** | **13.12M (0.31%)** | **\$4.27** | **74.9** |

RecursiveMAS trains the fewest parameters, uses the least memory, costs the least, and scores the highest.
The 0.31% checks out: 13.12M over the 4.21B base is 0.31% (reasoned). The interesting comparison is against
LoRA, which also trains a tiny adapter (0.37%) but on individual agents — it lands at 66.9 average accuracy
against RecursiveMAS's 74.9. Same order of trainable parameters, ~8 points apart, because LoRA improves each
agent in isolation while RecursiveLink optimizes the *collaboration* between them. That is the cleanest
evidence that the gain is coming from the system-level loop, not from the extra capacity of the links.

## Where it stands

The honest frame: RecursiveMAS is a well-constructed answer to a well-posed question. The looped-transformer
analogy is not a metaphor stretched over a MAS — it is implemented literally, with agents as layers, latent
states as the inter-layer signal, and a trained recurrence closing the loop, and the runtime-complexity
argument for why that is cheaper than text is concrete and falsifiable. The efficiency numbers scale the way
the theory predicts, and the cost table against LoRA is a genuinely convincing isolation of the effect.

The caveats are equally concrete. Everything is demonstrated on sub-5B open models — the paper's own
scaling plots are "sub-1.5B agents" — so whether system-level latent recursion still pays once the agents
are frontier-scale is untested here. The latent channel is unmonitorable, which is a real regression for
anyone who needs to audit what agents tell each other. And the flashiest accuracy rows lean on Pass@10. None
of that undoes the core result, which is modest and real: you can scale a team of frozen models by training
0.31% of their weights to pass thoughts instead of text, and recursion depth becomes a new axis you can turn
up for more accuracy and *less* cost at the same time. That combination — [unlike most ways of scaling
agents](/articles/jev-engineering-swarms), which trade accuracy for throughput or the reverse — is what
makes it worth reading.

For the broader looped-model context this sits in, the companion pieces are worth it: [what separates good
looped models from bad ones](/articles/looped-models-done-right), [whether looping buys reasoning under
matched compute](/articles/virtual-logic-depth), and [latent chain-of-thought at the single-model
level](/articles/lotus-latent-reasoning). The [heterogeneous-expert](/articles/mixture-of-experts-from-scratch)
and [attention](/articles/how-transformers-attention-works) background is assumed throughout.

---

*Built on ["Recursive Multi-Agent Systems"](https://arxiv.org/abs/2604.25917) (Zou, Pan, Qiu, Lu, Diao,
Jiang, Tong, Zhang, Buehler, He, Zou; NeurIPS 2026; arXiv 2604.25917v2). Code at
[RecursiveMAS/RecursiveMAS](https://github.com/RecursiveMAS/RecursiveMAS) (MIT); models and datasets at the
[RecursiveMAS Hugging Face org](https://huggingface.co/RecursiveMAS). All accuracy, speedup, token, and cost
figures are quoted from the paper's tables unless I label a recomputation as reasoned; the AIME figures are
Pass@10 as the paper reports them. The interactive diagram is an illustration of the mechanism — the moving
latents and the per-round readouts are drawn from the paper's reported numbers, not re-measured.*
