2026-10-02 · 15 min · llm · agents · multi-agent · looped-transformers · recurrent-depth · latent-reasoning · inference-optimization
A 1:48 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.
› transcript
Hi, I'm Ines! Let me show you how a whole team of agents can be scaled by looping them like one transformer. Not a bigger model, not more agents. Scale the collaboration itself. Loop the team in latent space and let it refine. The usual way, agents talk in text. Every hand-off decodes to words and then reads them back. This system keeps the hidden states instead, and trains the whole team as one. Three agents become the layers of one network. Here a planner, a solver, and a critic. Inside each agent an inner link turns its last hidden state into its next input. Latent thoughts, never written as words. An outer link carries those latent states to the next agent, even when the models have different sizes. The last agent loops back to the first. Round after round the team refines, and only the final round decodes to text. The link itself is tiny. A residual carries the embedding through, and a two-layer network learns only the gap between the spaces. A residual carries the embedding through. A two-layer net learns only the gap. Deeper recursion cuts more tokens. About thirty-five percent at one round, up to seventy-six at three, because the middle steps are never re-decoded. 13.12M trainable parameters, 0.31% of the weights. So the whole team trains as one. Agents as layers, latent links instead of text, and recursion as a new axis to scale. Agents become layers. Latent links, not text. Loop for more, at lower cost. Every source is in the full article. I'm Ines. Bye!
A looped transformer buys reasoning depth by reusing the same layers over and over instead of stacking new ones. You run the block, feed its output back as input, run it again — recurrent depth in latent space, no new parameters. LOTUS does the single-model version of this for chain-of-thought: it reasons in hidden states across looped passes rather than writing every step out as tokens.
RecursiveMAS asks the obvious next question. A multi-agent system (MAS) is already a stack of LLM calls — a Planner, then a Solver, then a Critic — wired in sequence. What if you treat that whole pipeline as one looped transformer, where each agent is a layer, the states passed between agents are latent vectors instead of text, and the last agent's output loops back into the first for another round? Only the final round decodes to words. Everything in between stays in embedding space.
That is the entire idea, and it is a genuinely different lever from the two that dominate agent research: not a bigger model, not more agents, but deeper connected collaboration over the same team. The paper ("Recursive Multi-Agent Systems", Zou et al., NeurIPS 2026) builds it on small open models, freezes all of them, and trains a 13-million-parameter connector that is 0.31% of the weights. The headline claims are +8.3% average accuracy over the strongest baseline, a 1.2×–2.4× end-to-end speedup, and a 34.6%–75.6% token reduction. Below I walk the mechanism from first principles, then check each of those numbers against the paper's tables — including one metric choice that inflates the most eye-catching rows.
- license
- MIT
- branch
- main
- tests
- none found
- source
- 642.0 kB
- commit date
- 2026-09-28
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-02 at cbfcaab — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow
shallow clone: counts describe the pinned tree, not the history
The usual way: agents talk in text
A standard MAS is a group of LLMs that pass natural language to each other. The Planner writes a plan in words; the Solver reads those words, re-tokenizes them, and writes a solution; the Critic reads that and writes feedback. Every hand-off is a full round trip through the vocabulary: decode hidden states to discrete tokens, then embed those tokens back into hidden states at the next agent.
Two costs come out of this. The obvious one is tokens — every intermediate message is generated one token at a time, which is where inference latency lives, and a recursive MAS that loops the whole team re-pays that cost every round. The subtler one is information loss: decoding to a single token sequence collapses a rich distribution over the vocabulary into one sampled path, and the next agent only ever sees the path, not the distribution behind it.
RecursiveMAS keeps the collaboration structure — the same Planner/Solver/Critic topology — but replaces the text channel between agents with a latent one. Scrub the recursion depth and flip the channel to see the two effects that produces:
Each agent is a layer of one looped network. The inner link feeds an agent's last-layer embedding back as its next input embedding — latent thoughts, no decoding. The outer link carries those latent states to the next agent, across different hidden widths. After the A₃ the latents loop back to A₁ — one recursion round — and only the final round decodes to text. Flip to the text channel and every hop pays to re-decode and re-encode; that is the baseline the speedup and token numbers are measured against.
Agents as layers: the inner and outer links
To make "each agent is a layer" literal, you need a way to turn one agent's hidden states into the next agent's inputs without going through text. That connector is RecursiveLink, and it comes in two flavours for the two kinds of hand-off in the system.
The first hand-off is within an agent. An ordinary LLM generates autoregressively: produce a token, embed it, feed it back in for the next step. RecursiveMAS wants the agent to keep thinking without emitting a token, so it takes the agent's last-layer embedding and maps it back into the input-layer embedding space for the next forward step. The paper calls the result an ongoing latent thought. That map is the inner link :
where are linear layers and is GELU. It is a two-layer MLP with a residual connection — tiny. The agent runs this for steps, producing a short continuous sequence of latent thoughts instead of decoded tokens.
The second hand-off is across agents, and heterogeneous agents do not share a hidden width — a 1.7B Planner and a 1.5B Solver have different . The outer link handles that with one extra linear layer on the residual branch to project the source width to the target width:

Why the residual term in both? Because the residual carries the original embedding through largely unchanged, so the MLP only has to learn the difference between the two embedding spaces, not the whole projection from scratch. That is the same reason residual connections stabilize deep nets, applied to a cross-model adapter: the link aligns distributions rather than relearning semantics, which the authors argue makes training more stable. The branches are small enough that the two links across the whole system come to just 13.12M trainable parameters (reported, Table 5).
Closing the loop
Chain the agents and you get the full architecture. Agent generates its latent thoughts through its inner link, the outer link projects them into 's input space, conditions on both its own prompt and 's transferred latents, and so on down the line. When the last agent finishes, its latent output — the system's current latent answer — is passed back to through the inner-outer link, closing a recurrent loop over the whole team. Each new round lets every agent reflect on the previous round's system output and refine. Only the last agent, in the final round , actually decodes to text.

The mapping to a looped transformer is now exact: agents are layers, a recursion round is one pass through the stack, and looping the last layer back to the first is the recurrence. The difference from a single-model looped transformer is only that the "layers" are whole frozen LLMs of different sizes, glued by trained links.
Why latent transfer is actually cheaper
It would be easy to assume the latent channel is faster just because "latents are smaller than text." The paper makes the sharper argument through runtime complexity. For a text-based recursive MAS with the same loop structure, each round costs (reported):
The term that matters is : generating intermediate tokens means projecting hidden states onto the full vocabulary (tens of thousands wide) at every step, for every agent, every round. RecursiveMAS removes that projection — it never decodes intermediate steps — and the term collapses to :
Since for these small models (a 150k-token vocabulary against a hidden size of a couple thousand), trading for is a real saving, and it compounds: it is paid times per round and once per recursion round. That is the whole reason the token count falls and the wall-clock speedup grows with recursion depth — deeper loops would have re-decoded more intermediate text, so skipping it saves more. The efficiency is structural, not a tuning artifact.
Training: freeze the models, train the wiring
Because the base LLMs are frozen, there is very little to optimize — only the links — but how you optimize them across a loop matters. The recipe is a two-stage inner-outer loop (reported).
The inner loop warms up each agent on its own, in parallel. For a training pair , it passes the ground-truth text through the agent's input embedding layer to get a target, then trains that agent's inner link so its generated latent thoughts align with that target under a cosine-similarity regression. In effect, each agent learns to "think" in latent form what it would otherwise have written in text, without the decode-and-re-encode step.
The outer loop then trains the system as one entity. It unrolls the full looped MAS for rounds and backpropagates through the entire recursive trace, so gradients flow across rounds and the outer links get shared credit — every link is updated for its contribution to the final answer, several rounds downstream. This is the part that makes it "whole-system co-optimization" rather than training agents in isolation.
There is a non-obvious reason the latent loop trains more stably than you might fear. The authors show that if you tried to do this through text with standard supervised fine-tuning, confident tokens (entropy with — i.e. a near-one-hot softmax) drive the gradient through the decoding step toward zero: gradient vanishing across the loop. RecursiveLink skips the softmax, and the paper proves its gradient norm stays close to 1 through the looped backpropagation (reported; a theoretical result under stated assumptions). Looping through discrete text is where recursive training breaks; looping through latents is where it keeps working.
Checking the headline numbers
The paper instantiates RecursiveMAS under four collaboration patterns — sequential, mixture, distillation, and deliberation — across nine benchmarks in math (MATH500, AIME2025, AIME2026), science and medicine (GPQA-Diamond, MedQA), code (LiveCodeBench-v6, MBPP+), and search QA (HotpotQA, Bamboogle). The sequential team I drew above is real and on Hugging Face: a Qwen3-1.7B Planner, a Qwen2.5-Math-1.5B Solver, and a Llama3.2-1B Critic for the "light" setting, scaled up to Gemma3-4B / Qwen3.5-4B / Llama3.2-3B.
Does accuracy scale with recursion depth? Yes, and this is the cleanest result. Averaged over the math/science/code tasks, RecursiveMAS beats the text-based recursive baseline by +3.4 points at , +6.0 at , and +7.2 at (reported, Table 2), with the relative margin widening from 8.1% to 20.2% as the loop deepens. Efficiency scales the same direction: the end-to-end speedup grows 1.2× → 1.9× → 2.4× and the token reduction grows 34.6% → 65.5% → 75.6% across (reported). More recursion is more advantage on both axes at once — the point of the stepper above.
The scaling has two knobs, not one: recursion depth at training time and at inference time are separate, and the paper varies both. Accuracy climbs toward the corner where both are large — training teaches the system to form refinement-ready latent states, and inference recursion then cashes that structure in for test-time gains. The same figure shows the framework is not tied to the sequential topology: it lifts the best standalone specialist under mixture, deliberation, and distillation patterns too.

The +8.3% whole-system claim. At , against a broad set of baselines (Table 3):
| Method | MATH500 | AIME2025 | AIME2026 | GPQA-D | LiveCodeBench | MedQA |
|---|---|---|---|---|---|---|
| Single Agent (Full-SFT) | 83.2 | 73.3 | 76.7 | 62.8 | 38.6 | 77.0 |
| Mixture-of-Agents | 79.8 | 60.0 | 63.3 | 47.6 | 27.0 | 57.5 |
| TextGrad | 84.9 | 73.3 | 76.7 | 62.5 | 39.8 | 77.2 |
| LoopLM | 84.6 | 66.7 | 63.3 | 48.1 | 24.9 | 56.4 |
| Recursive-TextMAS | 85.8 | 73.3 | 73.3 | 61.6 | 38.7 | 77.0 |
| RecursiveMAS | 88.0 | 86.7 | 86.7 | 66.2 | 42.9 | 79.3 |
RecursiveMAS wins every column — that part holds. But the +8.3% figure deserves a second look. The paper describes it as the average improvement "over the strongest baseline on each benchmark" across all nine benchmarks. On the six benchmarks shown here, if I take each column's strongest baseline and average the gap, I get +5.7 points (reasoned, from Table 3); RecursiveMAS's six-benchmark mean of ~75.0 sits about +5.9 points over the best overall baseline, TextGrad at ~69.1 (reasoned). The headline +8.3% is a nine-benchmark number, so the two search-QA sets and MBPP+ — not in this table — must carry larger margins to lift the average. The claim is the paper's (reported); I can confirm the direction and the per-column wins, but the +8.3 specifically is not reproducible from the comparison table alone.
The cost table. This is the efficiency argument that most interests me, because it is measured on the same scaled sequential setup, not estimated (reported, Table 5):
| Method | Peak GPU mem | Trainable params | Est. cost | Avg. acc |
|---|---|---|---|---|
| LoRA | 21.67 GB | 15.92M (0.37%) | $6.64 | 66.9 |
| Full-SFT | 41.40 GB | 4.21B (100%) | $9.67 | 68.6 |
| RecursiveMAS | 15.29 GB | 13.12M (0.31%) | $4.27 | 74.9 |
RecursiveMAS trains the fewest parameters, uses the least memory, costs the least, and scores the highest. The 0.31% checks out: 13.12M over the 4.21B base is 0.31% (reasoned). The interesting comparison is against LoRA, which also trains a tiny adapter (0.37%) but on individual agents — it lands at 66.9 average accuracy against RecursiveMAS's 74.9. Same order of trainable parameters, ~8 points apart, because LoRA improves each agent in isolation while RecursiveLink optimizes the collaboration between them. That is the cleanest evidence that the gain is coming from the system-level loop, not from the extra capacity of the links.
Where it stands
The honest frame: RecursiveMAS is a well-constructed answer to a well-posed question. The looped-transformer analogy is not a metaphor stretched over a MAS — it is implemented literally, with agents as layers, latent states as the inter-layer signal, and a trained recurrence closing the loop, and the runtime-complexity argument for why that is cheaper than text is concrete and falsifiable. The efficiency numbers scale the way the theory predicts, and the cost table against LoRA is a genuinely convincing isolation of the effect.
The caveats are equally concrete. Everything is demonstrated on sub-5B open models — the paper's own scaling plots are "sub-1.5B agents" — so whether system-level latent recursion still pays once the agents are frontier-scale is untested here. The latent channel is unmonitorable, which is a real regression for anyone who needs to audit what agents tell each other. And the flashiest accuracy rows lean on Pass@10. None of that undoes the core result, which is modest and real: you can scale a team of frozen models by training 0.31% of their weights to pass thoughts instead of text, and recursion depth becomes a new axis you can turn up for more accuracy and less cost at the same time. That combination — unlike most ways of scaling agents, which trade accuracy for throughput or the reverse — is what makes it worth reading.
For the broader looped-model context this sits in, the companion pieces are worth it: what separates good looped models from bad ones, whether looping buys reasoning under matched compute, and latent chain-of-thought at the single-model level. The heterogeneous-expert and attention background is assumed throughout.
Built on "Recursive Multi-Agent Systems" (Zou, Pan, Qiu, Lu, Diao, Jiang, Tong, Zhang, Buehler, He, Zou; NeurIPS 2026; arXiv 2604.25917v2). Code at RecursiveMAS/RecursiveMAS (MIT); models and datasets at the RecursiveMAS Hugging Face org. All accuracy, speedup, token, and cost figures are quoted from the paper's tables unless I label a recomputation as reasoned; the AIME figures are Pass@10 as the paper reports them. The interactive diagram is an illustration of the mechanism — the moving latents and the per-round readouts are drawn from the paper's reported numbers, not re-measured.