# LoopCD: a looped transformer's first loop is its own amateur

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/looped-contrastive-decoding
> date: 2026-10-06
> tags: looped-transformers, inference, reasoning

A looped transformer computes the same next token several times over. Each pass through its shared block produces a hidden state that the output head could decode, and ordinary decoding reads only the last one. The earlier states are computed, paid for, and dropped.

[Decoding Looped Transformers Better for (Almost) Free](https://arxiv.org/abs/2610.02185) (Liu, Zheng, Chen, Dilip, Bai, Jiao, Wang, Zhang; Apple, 1 October 2026) picks one of them back up. Its method, **LoopCD**, decodes the state after the *first* loop as well, treats it as a weak model's opinion, and pushes the final prediction further away from it. That is contrastive decoding, with the amateur model already inside the expert.

The paper reached me through [a Chinese-language post on X](https://x.com/mylifcc/status/2107291249921409470) that summarised it as: zero training cost, compute cut by 47%, math pass@1 up from 61% to 73%. The first is true. The other two are real numbers from two different experiments, and the post reads as if one run delivered both. Below: what a loop is, what contrastive decoding is, why a loop makes the amateur free, what the paper measured, and where the two headline numbers actually come from.

Every number here is labelled. **Reported** is the paper's figure, which I did not re-run (there is no code release I could find). **Reasoned** is my own arithmetic on reported numbers. **Measured** is something I computed from a file; in this article that is only the toy widget's arithmetic.

## A looped transformer, in one equation

A standard transformer buys depth with parameters: 48 layers means 48 sets of weights. A looped transformer stores a block once and runs it $R$ times, feeding each pass's output back in. The site's [architecture explainer](/architectures/looped-transformer) covers the design space; the short version is two layouts.

- **Loop the whole stack.** [Ouro](https://arxiv.org/abs/2510.25741) runs its full decoder 4 times: Ouro-1.4B is 24 layers looped to 96 layer applications, Ouro-2.6B is 48 layers looped to 192 (reported, Table A1). Only the final norm and vocabulary head sit after the loop.
- **Prelude, loop, coda.** [Huginn-0125](https://arxiv.org/abs/2502.05171) embeds the input through 2 prelude layers, iterates a 4-layer core 32 times by default, and decodes through a 2-layer coda: 132 layer applications at $R = 32$ (reported). Parcae and Looped-Qwen3 use the same sandwich. Looped-Qwen3 is a retrofit ([Training-Free Looped Transformers](https://arxiv.org/abs/2605.23872)): a frozen Qwen3-4B with a 4-layer middle window repeated as damped substeps, 15 layers before it and 17 after (reported).

<Figure
  src="https://ai.thesatyajit.com/articles/looped-contrastive-decoding/fig2.png"
  alt="Four layer diagrams. Ouro repeats its whole stack of layers. Huginn, Parcae and Looped-Qwen3 each have gray prelude layers, a blue shared block that repeats, and gray coda layers after it."
  caption="Which part loops. Ouro repeats its full stack; Huginn, Parcae and Looped-Qwen3 repeat a shared block (blue) between fixed prelude and coda layers (gray) (LoopCD paper, Figure 2)."
/>

The paper writes all four the same way. Let $h_0$ be the state entering the loop and $f$ the shared update:

$$
h_0 \xrightarrow{f} h_1 \xrightarrow{f} \cdots \xrightarrow{f} h_R, \qquad z_R = \mathrm{lm\_head}\big(\mathrm{coda}(h_R)\big)
$$

$z_R$ is the vector of next-token logits, and decoding samples from $p_R = \mathrm{softmax}(z_R)$. The point that makes this paper work: $h_1$ lives in the same space as $h_R$, because the same block wrote both. So $\mathrm{lm\_head}(\mathrm{coda}(h_1))$ is a legitimate next-token distribution too. It is just a worse one.

How much worse is reported in Figure 5 and Table A11. On the seven-benchmark multiple-choice mean, the prediction after the first loop trails the final one by 4.0 to 23.7 points across six models (reported). On ARC-Challenge, Huginn's first step scores 22.78% against 37.54% for its final prediction; Ouro-1.4B's first pass scores 38.65% against 60.41% (reported).

## Contrastive decoding, from the original

[Contrastive decoding](https://arxiv.org/abs/2210.15097) (Li et al., 2022) starts from an observation about failure modes. A large language model and a small one make many of the same mistakes: repetition, generic continuations, the most frequent token. Where the large model's preference differs from the small one's is where its extra capacity shows. So score each candidate token $v$ by the *difference*:

$$
\mathrm{score}(v) = \log p_{\text{expert}}(v) - \log p_{\text{amateur}}(v)
$$

A pure difference has an obvious bug: a token both models consider nearly impossible can win, because $10^{-9} / 10^{-12}$ is a big ratio. Li et al. guard against that with a **plausibility constraint**: only tokens with $p_{\text{expert}}(v) \ge \alpha \cdot \max_w p_{\text{expert}}(w)$ are eligible, with $\alpha = 0.1$.

[O'Brien and Lewis (2023)](https://arxiv.org/abs/2309.09117) rewrote it with a strength knob so it degrades gracefully to ordinary decoding, using expert logits $s_e$ and amateur logits $s_a$:

$$
\mathrm{score}(v) = (1 + \beta)\, s_e(v) - \beta\, s_a(v)
$$

$\beta = 0$ is the expert alone; they used $\beta = 0.5$ and kept $\alpha = 0.1$, and reported gains on GSM8K and HellaSwag with LLaMA-65B as the expert. The practical problem with all of this is the amateur. You need a second model that shares the tokenizer, makes the same kind of mistakes, and is cheap enough to run alongside the expert at every step. [DoLa](https://arxiv.org/abs/2309.03883) dodged it by using an early layer of the same model as the amateur, which needs that early layer to be decodable.

## LoopCD: the amateur was already computed

A looped transformer supplies an aligned amateur for nothing. $h_1$ is the same model, on the same prefix, after one loop's worth of compute instead of $R$. LoopCD-Logits decodes both and extrapolates:

$$
z_1 = \mathrm{lm\_head}(\mathrm{coda}(h_1)), \qquad z' = z_R + \omega\,(z_R - z_1), \quad \omega \ge 0
$$

Rearranged, $z' = (1 + \omega) z_R - \omega z_1$: exactly O'Brien and Lewis's form with $\beta = \omega$, the final loop as expert and the first loop as amateur. The paper's Appendix A.3 states the probability form, $p'(v) \propto p_R(v)\,\big(p_R(v)/p_1(v)\big)^{\omega}$. Read it as: keep the final distribution, then multiply each token by how much the loops raised it, to the power $\omega$.

There is no plausibility mask in the paper's equations (Appendix A.3 says its identity describes the logits "before any temperature or truncation used by an evaluation protocol"). Its guard is a different one: an **adaptive strength** that sets $\omega$ per token from how close the final prediction's top two candidates are,

$$
\omega = \omega_{\max}\big[1 - (p_{R,(1)} - p_{R,(2)})\big]
$$

so a confident token gets almost no push and a coin-flip gets the full $\omega_{\max}$.

<Figure
  src="https://ai.thesatyajit.com/articles/looped-contrastive-decoding/fig3.png"
  alt="Two pipelines. Top, LoopCD-Logits: h_R and h_1 each pass through the output layers, giving z_R and z_1, which are combined with weight omega into guided logits. Bottom, LoopCD-Hidden: h_R and h_1 are combined with weight omega first, and the result passes through the output layers once."
  caption="The two places the contrast can go. LoopCD-Logits decodes both states and combines logits (two output passes); LoopCD-Hidden combines the states and decodes once (LoopCD paper, Figure 3)."
/>

The second form, **LoopCD-Hidden**, applies the same extrapolation to the states before the coda, $h' = h_R + \omega (h_R - h_1)$, and decodes once. It costs no extra output pass. It is not the same function: the coda is nonlinear, so the probability identity above does not hold for it, and the adaptive rule needs $p_R$, which LoopCD-Hidden never computes. The paper runs LoopCD-Hidden at fixed strength only.

The widget below runs LoopCD-Logits on six invented tokens. The final loop slightly prefers "12" (logit 2.3) over "15" (1.9); the first loop preferred "12" by much more (2.5 against 0.8). Between loop one and loop $R$, "15" gained 1.1 logits and "12" lost 0.2, so the contrast favours "15", and once $\omega$ passes about 0.31 the argmax flips (0.4 logit gap, divided by 1.3 of relative motion). A junk token, "banana", shows why Li et al. needed $\alpha$: the first loop gave it a logit of −4.2 and the final loop −0.6, so its contrast of +3.6 is the largest in the vocabulary. Push $\omega$ past 1.0 and it wins. An $\alpha$ of 0.06 or more masks it; the adaptive rule only delays it. These logits are illustrative; the arithmetic is the paper's Eq. 3 and Eq. 4.

<ContrastLab />

### What the paper's own example looks like

The paper traces one ARC-Challenge question through Ouro-1.4B (Table A13, reported): which natural disaster leaves a narrow path of destruction through a forest. Per-character log-likelihoods:

| option | after the first pass | final | guided, $\omega = 0.5$ |
|---|---|---|---|
| a flood | −1.500 | −0.770 | −0.934 |
| a tornado (correct) | −1.113 | −0.209 | **−0.170** |
| a hurricane | −0.969 | −0.258 | −0.247 |
| an earthquake | −0.525 | **−0.201** | −0.199 |

The final pass picks "earthquake" by under 0.01. Across the passes, "tornado" rose by 0.90 and "earthquake" by 0.32; continuing that motion puts "tornado" first. The guided column is not $z_R + 0.5(z_R - z_1)$ applied to these four numbers, because guidance acts on every token of every option's continuation and each token is renormalised over the whole vocabulary; the option score is the sum afterwards. Same idea, one level down.

## Why the first loop, and not a later one

The intuitive choice for an amateur might be the loop just before the end: nearly as good, so a gentler, safer contrast. The paper's sweep says the opposite, and the reason is useful. What the reference supplies is a *direction*, and a late loop that has already converged on the final answer has no direction left to give.

Table A11 (reported) lists, for each reference loop $k$, its own accuracy, how often it picks a different answer from the final prediction, and the LoopCD-Logits gain at $\omega = 0.5$. On ARC-Challenge, Huginn's first step disagrees with the final prediction on 51.3% of questions and gives +3.07 points; its sixteenth disagrees on 7.3% and gives +0.17. Parcae-1.3B's sixth step is *more* accurate alone than the final prediction (41.04% against 40.36%) and gives +1.02, against +3.50 from its first. Huginn's second step is as weak as its first (22.10% against 22.78%) yet gives about half the gain, because it already agrees more with the end. Gain tracks disagreement, not competence.

<ReferencePicker />

One exception, and it is instructive. Huginn initialises its recurrent state with Gaussian noise. Its first step mostly removes that noise: the paper reports the state's norm stays at 76.4 within a spread of 0.05 across all 32 steps, while the path it walks is 6.0 times longer than its net displacement (reported, Appendix E.1). The coda projects the noise away, so $z_1$ is a fine logit reference. But LoopCD-Hidden contrasts the raw states and amplifies that noise: early states cost 0.58 points on Huginn, and the hidden-state gain peaks at the sixth step with +0.83 (reported). So the paper's Huginn LoopCD-Hidden runs use $h_6$ or $h_7$. Parcae and Looped-Qwen3 start from deterministic states and use $h_1$ in both forms.

Layers *inside* a pass are not a substitute. A DoLa-style reference taken from a mid-pass layer of Ouro dropped ARC-Challenge accuracy by up to 20.6 points on Ouro-1.4B and 24.9 on Ouro-2.6B (reported, Appendix E.3); within a pass, the decoded distribution swings to 0.57 to 0.84 bits of Jensen-Shannon divergence from the final one. In the paper's words, looped architectures "do not form trained exit points inside a pass"; the end of a completed loop is the natural one.

## What the contrast actually does: re-rank close calls

The paper splits the logit contrast $z_R - z_1$ into a part parallel to $z_R$ and a part orthogonal to it. The parallel part scales every logit by the same factor: a temperature change that sharpens the distribution without reordering it. The orthogonal part reorders. On HellaSwag, the orthogonal part alone gives +1.72 points on Ouro-1.4B at $\omega = 0.5$ against +0.76 for the full update (reported). Sharpening is not where the gain lives.

<Figure
  src="https://ai.thesatyajit.com/articles/looped-contrastive-decoding/fig8.png"
  alt="Three panels. (a) A vector diagram splitting the contrast into a component parallel to the final logits and an orthogonal component. (b) HellaSwag accuracy change for two Ouro models as guidance strength grows, re-ranking part alone above the full update. (c) ARC-Challenge accuracy change in five groups of questions by confidence, with large gains in the least confident group and near zero in the most confident."
  caption="The re-ranking part carries the gain, and the gain lands on uncertain decisions: grouped by the unguided prediction's confidence, the least confident fifth of ARC-Challenge questions gains the most (LoopCD paper, Figure 8)."
/>

A re-ranking can only flip a decision when its push exceeds the gap between the top two options. So the gain concentrates where the model is undecided. Splitting ARC-Challenge into fifths by that gap, guidance at $\omega = 0.5$ adds +6.4 to +13.3 points on the least confident fifth across four models and at most +0.4 on the most confident (reported). Only 5.8% to 14.9% of ARC-Challenge answers change at all (reported). This is a tie-breaker that continues the direction the loops were already moving in. It adds no knowledge; if the loops were heading the wrong way, it follows them.

## The results, at full depth

**Mathematical reasoning (Ouro-Thinking, sixteen samples per problem, Table 1a, reported).** Ouro-2.6B-Thinking with adaptive LoopCD-Logits ($\omega_{\max} = 1.5$):

| benchmark | baseline pass@1 | LoopCD pass@1 | change |
|---|---|---|---|
| AIME 2024 | 61.88 | 73.33 | +11.45 |
| AIME 2025 | 49.58 | 56.88 | +7.30 |
| OlympiadBench | 64.05 | 67.29 | +3.24 |

Mean pass@1 gains across the two Ouro-Thinking models and both rules run from 5.27 to 7.33 points; pass@10 from 2.53 to 4.12 (reported). This is the 61% to 73% in the post. It is at the model's full 4 loops, with an extra output pass, so it costs slightly *more* compute than the baseline, not less.

AIME is a small set: 30 problems per year, each estimated from 16 samples (reported, Appendix B.2). An 11.45-point pass@1 gain is about 3.4 problems' worth of expected per-sample success (reasoned: 0.1145 × 30). The paper reports no seeds, confidence intervals or sampling temperature that I could find, and its hyperparameters came from "screening sweeps prior to full-suite evaluations" without saying whether those sweeps used held-out data. I would treat the AIME number as the largest point in a consistent trend, not as a precise effect size.

**Code (Table 1b, Table 3b, reported).** The biggest code gain is LoopCD-Hidden on Huginn at $R = 32$: HumanEval pass@1 from 22.56% to 31.71%, with every code column improving at both depths. Ouro and Looped-Qwen3 gain 0.37 to 2.54 points on the four-column mean. Not every cell moves up: Huginn at $R = 32$ with the adaptive logit rule loses 0.79 on MBPP base.

**Multiple choice (Table 2, reported).** Every configuration's seven-benchmark mean improves under LoopCD-Logits, by +0.29 to +1.59 points. WinoGrande gets worse in most rows. Generation outside math and code is mixed: GSM8K improves in six of ten rows and drops by up to 1.29 points (Huginn, $R = 32$, fixed); Looped-Qwen3's AIME pass@1 *falls* under guidance (64.79 to 61.88 on AIME 2024, fixed) while its pass@10 rises (reported, Tables A5 and A6).

<Figure
  src="https://ai.thesatyajit.com/articles/looped-contrastive-decoding/fig1.png"
  alt="Three panels. (a) A row of shared blocks producing h_1 through h_R; h_R and h_1 each go through output layers, and the final and earlier predictions combine as final plus omega times final minus earlier. (b) Horizontal bars of baseline accuracy with LoopCD's gain added, for eight model and benchmark pairs, largest +11.45 on AIME 2024 with Ouro-2.6B Thinking. (c) Seven-benchmark mean against forward TFLOPs on a log axis, with arrows from the unguided full-depth model to LoopCD at half the iterations, labelled -22% to -46%."
  caption="The paper's overview: an earlier loop guides the final one; the full-depth gains (b); and half the iterations with LoopCD against the unguided full-depth model (c). Panels b and c are different experiments (LoopCD paper, Figure 1)."
/>

## Where the compute saving comes from

The "47%" is from a different question. If guidance adds accuracy at full depth, can it replace some of the loops? The paper halves the iteration count (Huginn 32 to 16, Parcae 8 to 4, Looped-Qwen3 8 to 4 substeps), applies LoopCD at that halved depth, and compares against the unguided model at full depth.

Halving the loops alone costs 0.17 to 1.29 points on the seven-benchmark mean; LoopCD at half depth adds back 0.75 to 1.29 and matches or beats the full-depth baseline in all six settings (reported). Huginn at 16 of its 32 iterations beats its own full-depth baseline by 1.02 points with LoopCD-Logits and 0.51 with LoopCD-Hidden; Looped-Qwen3 ties it at +0.00 (reported).

<Figure
  src="https://ai.thesatyajit.com/articles/looped-contrastive-decoding/fig4.png"
  alt="Two panels for six settings. (a) Change in seven-benchmark mean from the unguided full-depth model: unguided at half depth sits below zero, from -0.17 to -1.29, and LoopCD at half depth sits at or above zero, from +0.00 to +1.02. (b) Forward FLOPs at half depth as a share of the full pass, with savings of 46%, 48%, 22%, 27%, 39% and 24%."
  caption="Half the iterations with LoopCD against the unguided model at full depth: accuracy (a) and forward FLOPs, guidance included (b). Multiple-choice benchmarks only (LoopCD paper, Figure 4)."
/>

Compute is analytic forward FLOPs for a 512-token prefill, guidance included (reported, Table A9):

| setting | iterations | FLOPs vs full pass | removed |
|---|---|---|---|
| Huginn-0125, logits | 32 → 16 | 0.54 | 46.0% |
| Huginn-0125, hidden states | 32 → 16 | 0.52 | 48.2% |
| Parcae-370M, logits | 8 → 4 | 0.78 | 22.5% |
| Parcae-1.3B, logits | 8 → 4 | 0.73 | 27.3% |
| Parcae-1.3B, hidden states | 8 → 4 | 0.61 | 39.2% |
| Looped-Qwen3, hidden states | 8 → 4 | 0.76 | 23.6% |

No row says 47%. The post's figure sits between Huginn's 46.0% and 48.2%, the two best cases; the paper's range is 22.5% to 48.2%. The ceiling is set by architecture: halving the loops can only remove the loop's share of the pass. Huginn spends most of its FLOPs in the loop, so it approaches half. Looped-Qwen3 has 32 non-looped layers around a 4-layer window, so it cannot.

Three caveats the post dropped, all from the paper itself:

1. **The half-depth runs are multiple-choice only** (reported, Appendix D.1). There is no half-depth AIME number. The 73.33% and the 48.2% never occur in the same run, on the same model, or on the same kind of task.
2. **Ouro is not in the half-depth study at all.** The model behind the AIME number has 4 loops, and the paper does not test it at 2.
3. **FLOPs, not wall-clock.** The paper says its counts "reflect theoretical arithmetic workload rather than end-to-end wall-clock speedups" (reported). Loops run sequentially, so halving them should cut decode latency roughly in proportion on a loop-dominated model (reasoned), but nobody measured it.

## What "(almost) free" costs

At full depth, LoopCD-Logits pays for one extra pass through whatever sits after the loop (reported, Table A7):

| model | layers after the loop | LoopCD-Logits FLOPs |
|---|---|---|
| Ouro-1.4B | 0 | 1.019× |
| Ouro-2.6B | 0 | 1.010× |
| Huginn-0125, $R = 32$ | 2 | 1.022× |
| Parcae-1.3B | 8 | 1.119× |
| Looped-Qwen3 | 17 | 1.306× |

For Ouro, "almost free" means 1% to 2%: a second norm and vocabulary projection. For Looped-Qwen3 it means 30.6% more compute, because the reference has to run the frozen 17-layer tail. That is why its half-depth result uses LoopCD-Hidden: with the logit form, the readout overhead would exceed the loop savings (the paper's reading of Figure A1). LoopCD-Hidden costs 1.000× everywhere, but a contrast injected before a deep coda gets damped: on Parcae-1.3B's 8-layer coda it keeps roughly 60% of the logit form's multiple-choice gain (0.85 against 1.38 points, reported).

Memory is the cost the paper does not tabulate. LoopCD needs $h_1$ (or $h_6$) to still exist when $h_R$ is done: one extra hidden vector per position being decoded (reasoned). At decode time that is small next to the KV cache. In a fused serving kernel that discards intermediate loop states, it is a plumbing change, not a free switch.

## The strength knob is narrow for generation

Multiple-choice scoring tolerates a wide band of fixed $\omega$, peaking near 0.5 (reported, Figure 9). Generation does not: an early token pushed the wrong way changes every token after it. The paper picks $\omega$ of 0.2 to 0.3 for generation, and beyond about 0.6 the curves fall steeply; Ouro-1.4B loses 8 points on GSM8K at $\omega = 0.8$ and 20 at $\omega = 1.0$ (reported, Appendix G.1). The adaptive rule widens the safe range, staying positive for caps from 0.5 to 1.0 where fixed strength overshoots (reported). Read the 73.33% with that in mind: it is the adaptive rule at a cap of 1.5 on a model trained, unlike Huginn and Parcae, to make its intermediate loops decodable.

## Where it sits among the site's looped-model coverage

This paper is the first in this cluster to ask what to do with a looped model at inference rather than how to train one. [Towards looped models done right](/articles/looped-models-done-right) separates Ouro-style and Huginn-style design choices; the random-init axis it studies is exactly why Huginn's $h_1$ fails as a hidden-state reference here. [SMELT](/articles/looped-transformers-matched-compute) and [LOOM](/articles/loom-looped-moe) are about whether more loops pay for their FLOPs at training time, and [virtual logic depth](/articles/virtual-logic-depth) about what looping buys at all. LoopCD's half-depth result is a different trade: it does not make loops cheaper, it makes the half of them you kept work harder. (Unrelated despite the name: the site's [Contrastive Language Model](/articles/contrastive-language-model) piece is about contrastive embeddings, not contrastive decoding.)

## What I'd take from it

- **The amateur problem disappears for looped models.** Contrastive decoding always needed a weak model that shares the expert's vocabulary and failure modes. A loop's first pass is that model, and it is already computed (reported mechanism; the contrast form is O'Brien and Lewis's with $\beta = \omega$).
- **Pick the reference by disagreement, not accuracy.** The first completed loop wins for logits; noise-initialised models need a post-burn-in state for the hidden form (reported).
- **The 73.33% and the 48.2% are separate claims.** One is full-depth math with a slightly larger compute bill; the other is half-depth multiple-choice. Neither has a wall-clock measurement or error bars (reported; reasoned on what is absent).
- **It is a tie-breaker.** It moves 5.8% to 14.9% of ARC-Challenge answers, almost all of them close calls (reported). It cannot rescue a loop that is converging on the wrong answer.

I have not run it. The method is about ten lines on top of any looped model that exposes its per-loop states, and the [Ouro checkpoints](https://huggingface.co/ByteDance/Ouro-2.6B-Thinking) are public, so the AIME claim is checkable by anyone with a GPU and patience for 480 sampled solutions per AIME year (30 problems, 16 samples each).

<ModelCard repo="ByteDance/Ouro-2.6B-Thinking" />
