~/satyajit

RecursiveMAS: a multi-agent system folded into one looped transformer

mdjsonmcp

2026-10-02 · 15 min · llm · agents · multi-agent · looped-transformers · recurrent-depth · latent-reasoning · inference-optimization

A 1:48 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.

› transcript

Hi, I'm Ines! Let me show you how a whole team of agents can be scaled by looping them like one transformer. Not a bigger model, not more agents. Scale the collaboration itself. Loop the team in latent space and let it refine. The usual way, agents talk in text. Every hand-off decodes to words and then reads them back. This system keeps the hidden states instead, and trains the whole team as one. Three agents become the layers of one network. Here a planner, a solver, and a critic. Inside each agent an inner link turns its last hidden state into its next input. Latent thoughts, never written as words. An outer link carries those latent states to the next agent, even when the models have different sizes. The last agent loops back to the first. Round after round the team refines, and only the final round decodes to text. The link itself is tiny. A residual carries the embedding through, and a two-layer network learns only the gap between the spaces. A residual carries the embedding through. A two-layer net learns only the gap. Deeper recursion cuts more tokens. About thirty-five percent at one round, up to seventy-six at three, because the middle steps are never re-decoded. 13.12M trainable parameters, 0.31% of the weights. So the whole team trains as one. Agents as layers, latent links instead of text, and recursion as a new axis to scale. Agents become layers. Latent links, not text. Loop for more, at lower cost. Every source is in the full article. I'm Ines. Bye!

A looped transformer buys reasoning depth by reusing the same layers over and over instead of stacking new ones. You run the block, feed its output back as input, run it again — recurrent depth in latent space, no new parameters. LOTUS does the single-model version of this for chain-of-thought: it reasons in hidden states across RR looped passes rather than writing every step out as tokens.

RecursiveMAS asks the obvious next question. A multi-agent system (MAS) is already a stack of LLM calls — a Planner, then a Solver, then a Critic — wired in sequence. What if you treat that whole pipeline as one looped transformer, where each agent is a layer, the states passed between agents are latent vectors instead of text, and the last agent's output loops back into the first for another round? Only the final round decodes to words. Everything in between stays in embedding space.

That is the entire idea, and it is a genuinely different lever from the two that dominate agent research: not a bigger model, not more agents, but deeper connected collaboration over the same team. The paper ("Recursive Multi-Agent Systems", Zou et al., NeurIPS 2026) builds it on small open models, freezes all of them, and trains a 13-million-parameter connector that is 0.31% of the weights. The headline claims are +8.3% average accuracy over the strongest baseline, a 1.2×–2.4× end-to-end speedup, and a 34.6%–75.6% token reduction. Below I walk the mechanism from first principles, then check each of those numbers against the paper's tables — including one metric choice that inflates the most eye-catching rows.

RecursiveMAS/RecursiveMAS@cbfcaab · snapshot 2026-10-02
tracked files
43
license
MIT
branch
main
tests
none found
source
642.0 kB
commit date
2026-09-28
source by language
Python642.0 kB(29)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-02 at cbfcaab — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow

shallow clone: counts describe the pinned tree, not the history

The usual way: agents talk in text

A standard MAS is a group of LLMs that pass natural language to each other. The Planner writes a plan in words; the Solver reads those words, re-tokenizes them, and writes a solution; the Critic reads that and writes feedback. Every hand-off is a full round trip through the vocabulary: decode hidden states to discrete tokens, then embed those tokens back into hidden states at the next agent.

Two costs come out of this. The obvious one is tokens — every intermediate message is generated one token at a time, which is where inference latency lives, and a recursive MAS that loops the whole team re-pays that cost every round. The subtler one is information loss: decoding to a single token sequence collapses a rich distribution over the vocabulary into one sampled path, and the next agent only ever sees the path, not the distribution behind it.

RecursiveMAS keeps the collaboration structure — the same Planner/Solver/Critic topology — but replaces the text channel between agents with a latent one. Scrub the recursion depth and flip the channel to see the two effects that produces:

multi-agent system as a looped transformerlatent channel
recursion rounds123last agent → first agent, refineQpromptinner linkA₁ · PlannerQwen3-1.7Bfrozen · layer 1inner linkA₂ · SolverQwen2.5-Math-1.5Bfrozen · layer 2inner linkA₃ · CriticLlama3.2-1Bfrozen · layer 3outer linklatent stateouter linklatent statetext answerfinal round onlyrecurrence · last agent → first agent
channel
round r
2.4×
end-to-end speedup
−75.6%
token usage
+7.2
acc. vs text-MAS (pts)

Each agent is a layer of one looped network. The inner link feeds an agent's last-layer embedding back as its next input embedding — latent thoughts, no decoding. The outer link carries those latent states to the next agent, across different hidden widths. After the A₃ the latents loop back to A₁ — one recursion round — and only the final round decodes to text. Flip to the text channel and every hop pays to re-decode and re-encode; that is the baseline the speedup and token numbers are measured against.

To make "each agent is a layer" literal, you need a way to turn one agent's hidden states into the next agent's inputs without going through text. That connector is RecursiveLink, and it comes in two flavours for the two kinds of hand-off in the system.

The first hand-off is within an agent. An ordinary LLM generates autoregressively: produce a token, embed it, feed it back in for the next step. RecursiveMAS wants the agent to keep thinking without emitting a token, so it takes the agent's last-layer embedding hh and maps it back into the input-layer embedding space for the next forward step. The paper calls the result an ongoing latent thought. That map is the inner link Rin\mathcal{R}_{\text{in}}:

Rin(h)=h+W2 σ(W1h)\mathcal{R}_{\text{in}}(h) = h + W_2\,\sigma(W_1 h)

where W1,W2W_1, W_2 are linear layers and σ\sigma is GELU. It is a two-layer MLP with a residual connection — tiny. The agent runs this for mm steps, producing a short continuous sequence of latent thoughts [ht,ht+1,…,ht+m][h_t, h_{t+1}, \dots, h_{t+m}] instead of mm decoded tokens.

The second hand-off is across agents, and heterogeneous agents do not share a hidden width — a 1.7B Planner and a 1.5B Solver have different dhd_h. The outer link Rout\mathcal{R}_{\text{out}} handles that with one extra linear layer W3W_3 on the residual branch to project the source width to the target width:

Rout(h)=W3h+W2 σ(W1h)\mathcal{R}_{\text{out}}(h) = W_3 h + W_2\,\sigma(W_1 h)
Two side-by-side diagrams of the RecursiveLink module. Left (Inner): a Last-layer Emb feeds a stack of Linear, GELU, Linear, whose output is summed at a plus node with a residual connection carrying the Last-layer Emb directly, producing an Input-layer Emb for the same agent. Right (Outer): an Agent A_i Emb feeds the same Linear, GELU, Linear stack, but the residual branch now passes through an additional Linear block before the plus node, producing an Agent A_j Emb in the target agent's space.
RecursiveLink is a two-layer residual MLP. The inner link maps an agent's last-layer embedding back to its own input space; the outer link adds a linear W3 on the residual branch to map between agents of different hidden widths (RecursiveMAS, Figure 3).

Why the residual term in both? Because the residual carries the original embedding through largely unchanged, so the MLP only has to learn the difference between the two embedding spaces, not the whole projection from scratch. That is the same reason residual connections stabilize deep nets, applied to a cross-model adapter: the link aligns distributions rather than relearning semantics, which the authors argue makes training more stable. The branches are small enough that the two links across the whole system come to just 13.12M trainable parameters (reported, Table 5).

Closing the loop

Chain the agents and you get the full architecture. Agent A1A_1 generates its latent thoughts through its inner link, the outer link projects them into A2A_2's input space, A2A_2 conditions on both its own prompt and A1A_1's transferred latents, and so on down the line. When the last agent ANA_N finishes, its latent output — the system's current latent answer — is passed back to A1A_1 through the inner-outer link, closing a recurrent loop over the whole team. Each new round lets every agent reflect on the previous round's system output and refine. Only the last agent, in the final round nn, actually decodes to text.

The RecursiveMAS architecture as a left-to-right pipeline. Agent A1 produces a row of latent thoughts (last-layer embeddings h_t, h_{t+1}, ...) that are routed by an inner link back into its own input embeddings. Those latents are then passed through an outer link to Agent A2, whose inputs are the A2-aligned embeddings conditioned on A1's output; A2 runs its own inner link. This continues to Agent A_N, which decodes the final outputs in the last recursion round. A looping symbol marks that the last agent feeds back to the first for rounds before the last.
Each agent generates latent thoughts with its inner link, then transfers them to the next agent through the outer link. After the last agent, its latents feed back to the first, forming a recursive loop; only the final round decodes to text (RecursiveMAS, Figure 2).

The mapping to a looped transformer is now exact: agents are layers, a recursion round is one pass through the stack, and looping the last layer back to the first is the recurrence. The difference from a single-model looped transformer is only that the "layers" are whole frozen LLMs of different sizes, glued by trained links.

Why latent transfer is actually cheaper

It would be easy to assume the latent channel is faster just because "latents are smaller than text." The paper makes the sharper argument through runtime complexity. For a text-based recursive MAS with the same loop structure, each round costs (reported):

Θ ⁣(N( m∣V∣dh+(t+m)dh2+(t+m)2dh ))\Theta\!\big(N(\,m|V|d_h + (t{+}m)d_h^2 + (t{+}m)^2 d_h\,)\big)

The term that matters is m∣V∣dhm|V|d_h: generating mm intermediate tokens means projecting hidden states onto the full vocabulary ∣V∣|V| (tens of thousands wide) at every step, for every agent, every round. RecursiveMAS removes that projection — it never decodes intermediate steps — and the term collapses to mdh2m d_h^2:

Θ ⁣(N( mdh2+(t+m)dh2+(t+m)2dh ))\Theta\!\big(N(\,m d_h^2 + (t{+}m)d_h^2 + (t{+}m)^2 d_h\,)\big)

Since ∣V∣≫dh|V| \gg d_h for these small models (a 150k-token vocabulary against a hidden size of a couple thousand), trading m∣V∣dhm|V|d_h for mdh2m d_h^2 is a real saving, and it compounds: it is paid NN times per round and once per recursion round. That is the whole reason the token count falls and the wall-clock speedup grows with recursion depth — deeper loops would have re-decoded more intermediate text, so skipping it saves more. The efficiency is structural, not a tuning artifact.

Training: freeze the models, train the wiring

Because the base LLMs are frozen, there is very little to optimize — only the links — but how you optimize them across a loop matters. The recipe is a two-stage inner-outer loop (reported).

The inner loop warms up each agent on its own, in parallel. For a training pair (x,y)(x, y), it passes the ground-truth text yy through the agent's input embedding layer to get a target, then trains that agent's inner link so its generated latent thoughts align with that target under a cosine-similarity regression. In effect, each agent learns to "think" in latent form what it would otherwise have written in text, without the decode-and-re-encode step.

The outer loop then trains the system as one entity. It unrolls the full looped MAS for nn rounds and backpropagates through the entire recursive trace, so gradients flow across rounds and the outer links get shared credit — every link is updated for its contribution to the final answer, several rounds downstream. This is the part that makes it "whole-system co-optimization" rather than training agents in isolation.

There is a non-obvious reason the latent loop trains more stably than you might fear. The authors show that if you tried to do this through text with standard supervised fine-tuning, confident tokens (entropy ≤ϵ\le \epsilon with ϵ≪1\epsilon \ll 1 — i.e. a near-one-hot softmax) drive the gradient through the decoding step toward zero: gradient vanishing across the loop. RecursiveLink skips the softmax, and the paper proves its gradient norm stays close to 1 through the looped backpropagation (reported; a theoretical result under stated assumptions). Looping through discrete text is where recursive training breaks; looping through latents is where it keeps working.

Checking the headline numbers

The paper instantiates RecursiveMAS under four collaboration patterns — sequential, mixture, distillation, and deliberation — across nine benchmarks in math (MATH500, AIME2025, AIME2026), science and medicine (GPQA-Diamond, MedQA), code (LiveCodeBench-v6, MBPP+), and search QA (HotpotQA, Bamboogle). The sequential team I drew above is real and on Hugging Face: a Qwen3-1.7B Planner, a Qwen2.5-Math-1.5B Solver, and a Llama3.2-1B Critic for the "light" setting, scaled up to Gemma3-4B / Qwen3.5-4B / Llama3.2-3B.

Does accuracy scale with recursion depth? Yes, and this is the cleanest result. Averaged over the math/science/code tasks, RecursiveMAS beats the text-based recursive baseline by +3.4 points at r=1r=1, +6.0 at r=2r=2, and +7.2 at r=3r=3 (reported, Table 2), with the relative margin widening from 8.1% to 20.2% as the loop deepens. Efficiency scales the same direction: the end-to-end speedup grows 1.2× → 1.9× → 2.4× and the token reduction grows 34.6% → 65.5% → 75.6% across r=1,2,3r = 1, 2, 3 (reported). More recursion is more advantage on both axes at once — the point of the stepper above.

The scaling has two knobs, not one: recursion depth at training time and at inference time are separate, and the paper varies both. Accuracy climbs toward the corner where both are large — training teaches the system to form refinement-ready latent states, and inference recursion then cashes that structure in for test-time gains. The same figure shows the framework is not tied to the sequential topology: it lifts the best standalone specialist under mixture, deliberation, and distillation patterns too.

Top row: four heatmaps titled RecursiveMAS Scaling Law for MATH500, AIME2025, GPQA-D, and Code Gen, each a grid of training recursion round (x-axis, 1 to 4) against inference recursion round (y-axis, 1 to 4). Only the lower-right triangle is filled, and accuracy increases toward the top-right corner where both training and inference recursion are deep. Bottom row: three grouped bar charts titled Collaboration Patterns — Mixture-Style, Deliberation-Style, and Distillation-Style — where the dark RecursiveMAS bar is highest in almost every benchmark group against the standalone specialist agents, with per-benchmark end-to-end speedups around 1.4 to 1.6 times annotated on the distillation chart.
RecursiveMAS scales along two axes at once — training-time and inference-time recursion depth — with accuracy highest where both are deep (top), and the gains hold across mixture, deliberation, and distillation collaboration patterns, not just the sequential one (bottom) (RecursiveMAS, Figure 1).

The +8.3% whole-system claim. At r=3r=3, against a broad set of baselines (Table 3):

MethodMATH500AIME2025AIME2026GPQA-DLiveCodeBenchMedQA
Single Agent (Full-SFT)83.273.376.762.838.677.0
Mixture-of-Agents79.860.063.347.627.057.5
TextGrad84.973.376.762.539.877.2
LoopLM84.666.763.348.124.956.4
Recursive-TextMAS85.873.373.361.638.777.0
RecursiveMAS88.086.786.766.242.979.3

RecursiveMAS wins every column — that part holds. But the +8.3% figure deserves a second look. The paper describes it as the average improvement "over the strongest baseline on each benchmark" across all nine benchmarks. On the six benchmarks shown here, if I take each column's strongest baseline and average the gap, I get +5.7 points (reasoned, from Table 3); RecursiveMAS's six-benchmark mean of ~75.0 sits about +5.9 points over the best overall baseline, TextGrad at ~69.1 (reasoned). The headline +8.3% is a nine-benchmark number, so the two search-QA sets and MBPP+ — not in this table — must carry larger margins to lift the average. The claim is the paper's (reported); I can confirm the direction and the per-column wins, but the +8.3 specifically is not reproducible from the comparison table alone.

The cost table. This is the efficiency argument that most interests me, because it is measured on the same scaled sequential setup, not estimated (reported, Table 5):

MethodPeak GPU memTrainable paramsEst. costAvg. acc
LoRA21.67 GB15.92M (0.37%)$6.6466.9
Full-SFT41.40 GB4.21B (100%)$9.6768.6
RecursiveMAS15.29 GB13.12M (0.31%)$4.2774.9

RecursiveMAS trains the fewest parameters, uses the least memory, costs the least, and scores the highest. The 0.31% checks out: 13.12M over the 4.21B base is 0.31% (reasoned). The interesting comparison is against LoRA, which also trains a tiny adapter (0.37%) but on individual agents — it lands at 66.9 average accuracy against RecursiveMAS's 74.9. Same order of trainable parameters, ~8 points apart, because LoRA improves each agent in isolation while RecursiveLink optimizes the collaboration between them. That is the cleanest evidence that the gain is coming from the system-level loop, not from the extra capacity of the links.

Where it stands

The honest frame: RecursiveMAS is a well-constructed answer to a well-posed question. The looped-transformer analogy is not a metaphor stretched over a MAS — it is implemented literally, with agents as layers, latent states as the inter-layer signal, and a trained recurrence closing the loop, and the runtime-complexity argument for why that is cheaper than text is concrete and falsifiable. The efficiency numbers scale the way the theory predicts, and the cost table against LoRA is a genuinely convincing isolation of the effect.

The caveats are equally concrete. Everything is demonstrated on sub-5B open models — the paper's own scaling plots are "sub-1.5B agents" — so whether system-level latent recursion still pays once the agents are frontier-scale is untested here. The latent channel is unmonitorable, which is a real regression for anyone who needs to audit what agents tell each other. And the flashiest accuracy rows lean on Pass@10. None of that undoes the core result, which is modest and real: you can scale a team of frozen models by training 0.31% of their weights to pass thoughts instead of text, and recursion depth becomes a new axis you can turn up for more accuracy and less cost at the same time. That combination — unlike most ways of scaling agents, which trade accuracy for throughput or the reverse — is what makes it worth reading.

For the broader looped-model context this sits in, the companion pieces are worth it: what separates good looped models from bad ones, whether looping buys reasoning under matched compute, and latent chain-of-thought at the single-model level. The heterogeneous-expert and attention background is assumed throughout.


Built on "Recursive Multi-Agent Systems" (Zou, Pan, Qiu, Lu, Diao, Jiang, Tong, Zhang, Buehler, He, Zou; NeurIPS 2026; arXiv 2604.25917v2). Code at RecursiveMAS/RecursiveMAS (MIT); models and datasets at the RecursiveMAS Hugging Face org. All accuracy, speedup, token, and cost figures are quoted from the paper's tables unless I label a recomputation as reasoned; the AIME figures are Pass@10 as the paper reports them. The interactive diagram is an illustration of the mechanism — the moving latents and the per-round readouts are drawn from the paper's reported numbers, not re-measured.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "RecursiveMAS: a multi-agent system folded into one looped transformer", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026recursivemas,
  author = {Satyajit Ghana},
  title  = {RecursiveMAS: a multi-agent system folded into one looped transformer},
  url    = {https://ai.thesatyajit.com/articles/recursivemas},
  year   = {2026}
}
share