~/satyajit

LiFT: a looped DiT that keeps improving past the loops it was trained with

mdjsonmcp

2026-10-06 · 26 min · looped-transformers · flow-matching · diffusion · diffusion-transformers · image-generation · inference

Why read this

Notabletop 60%

The LiFT target derived as flow matching along depth, its FLOP accounting re-derived, and the 52% headline traced to a 250-step dense baseline.

  • Original analysis
  • A lasting reference
  • A new technique

Image & video generationAPI onlyResearch paper

How this was scored
Is it new?
2 of 3: A real new idea, method or capability
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
1 of 3: API-only, gated or restrictive licence
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
2 of 3: A reference for a year or more
Does it affect many?
1 of 3: A specialist community
Only here?
2 of 3: A teardown or measurement few others did

Score 62 of 100, ranked 202 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Earlier today this site ran a piece on LOOM and Looped-DiT, two papers about looping a model more than twice. Both push the useful loop count past two, and in both the loop count is something you fix at training time. Running a looped model deeper than it was trained is where they usually break, and the looped transformer page and the Ouro-vs-Huginn ablations show how many design choices sit underneath that.

Then Pedro Curvo and Mohammad Mahdi Derakhshani posted LiFT. One of its checkpoints was trained with 2 loops, and at 16 loops it gets better: FID 17.07 at two, 9.31 at sixteen, same weights, no retraining. Eight times the training depth, and the curve is still going down at the point where every other looped generator I've read about has turned up.

I expected a better normaliser, or a cleverer form of deep supervision. It is neither. LiFT changes one thing: what each loop is asked to predict. The new target turns out to be flow matching again, one level down, along depth instead of time. Once you see that, the extrapolation stops looking like luck.

The paper is arXiv 2610.05538 (University of Amsterdam and TNO). Its code, weights and W&B logs are listed as "Coming Soon". There is a Gradio demo running the L/2 checkpoint next to two dense baselines. So everything below is checked against the paper's equations, its appendix tables and its own FLOP formula, which I re-derived.

Two ways a flow model can spend compute

A flow-matching model is trained on straight lines. Draw noise x0∼N(0,I)x_0 \sim \mathcal{N}(0, I) and a data point x1x_1, and put a point on the segment between them:

xt=(1−t) x0+t x1,u⋆=dxtdt=x1−x0,t∈[0,1].x_t = (1-t)\,x_0 + t\,x_1, \qquad u^\star = \frac{\mathrm{d}x_t}{\mathrm{d}t} = x_1 - x_0, \qquad t \in [0,1].

The network vθ(xt,t)v_\theta(x_t, t) regresses u⋆u^\star with a plain MSE. Because many pairs pass through the same (xt,t)(x_t, t), the best it can do is their average, the marginal velocity v⋆(x,t)=E[u⋆∣xt=x,t]v^\star(x, t) = \mathbb{E}[u^\star \mid x_t = x, t]. (The CSFM write-up goes deeper on why that average matters.) Sampling integrates the learned field from noise to data. With Euler and TT steps, each step costs one network call:

xt+Δt=xt+Δt  vθ(xt,t),Δt=1/T.x_{t+\Delta t} = x_t + \Delta t \; v_\theta(x_t, t), \qquad \Delta t = 1/T.

So the usual inference knob is TT. More steps follow the field more faithfully. But an Euler step can only be as good as the velocity it is handed, and once integration error stops being the bottleneck, more steps buy almost nothing. The paper's own dense L/2 shows this: 16.94 at 50 steps, 16.17 at 100, 15.74 at 250. Five times the compute for about one FID point.

A looped network adds a second knob. Inside each call, run the same core KK times to get a better velocity estimate. The cost model separates cleanly:

Cinfer=T⋅Cforward(K).C_{\mathrm{infer}} = T \cdot C_{\mathrm{forward}}(K).

The thread puts it as "one axis is time; the other is depth". The paper's motto is "loop in depth, flow in time". The idea is easy to state. The hard part is training a core that keeps improving as KK grows.

What the network looks like

Four-panel method diagram. (a) Flow in time: x0 (noise) to x_t to x1 (data), with one dashed sampler step adding delta-t times u_K. (b) Loop in depth: x_t and c enter Prelude P_theta producing h0; h0 feeds three red Core R_theta boxes in sequence, each taking a depth coordinate s1, s2, s_K=1 from below, with h0 reinjected through plus nodes before the second and third cores; after each state a coda D_theta reads out a prediction u0, u1, u2 and finally u_K. (c) Training targets: in velocity space, a straight line from b = sg(u0) at s=0 to u-star at s=1, with hollow target points u-bar at s1 and s2; predictions u1, u2, u_K sit above the line, joined to their targets by dotted MSE lines. (d) Depth coordinates: training with K_train=3 draws s1, s2 uniformly and sorts them; inference with K_inf=4 uses 0, 1/4, 2/4, 3/4, 1.
LiFT expands one sampler step (a) into a loop over a shared core (b). Every state is read out to a velocity; in training each readout is pulled toward its own point on the line from the stopped-gradient first guess to the target (c), at depth coordinates drawn at random in training and spaced uniformly at inference (d) (LiFT paper, Figure 2).

The backbone is an ordinary DiT, split three ways. If you want the DiT itself from the ground up, the Diffusion Transformer architecture page covers patches, adaLN-Zero and the rest. The split, in the paper's Eq. 9:

h0=Pθ(xt,c),hk=Rθ(RMSNorm(hk−1)+h0,  c+es(sk)),k=1,…,K,uk=Dθ(hk,c),k=0,…,K.\begin{aligned} h_0 &= P_\theta(x_t, c), \\ h_k &= R_\theta\big(\mathrm{RMSNorm}(h_{k-1}) + h_0,\; c + e_s(s_k)\big), \qquad k = 1, \dots, K, \\ u_k &= \mathcal{D}_\theta(h_k, c), \qquad k = 0, \dots, K. \end{aligned}

The prelude PθP_\theta is four ordinary DiT blocks that embed the noisy latent once per sampling step. The core RθR_\theta is a stack of LRL_R blocks applied KK times with the same weights. Each pass normalises the previous state and adds h0h_0 back in, so the input never fades out of the recurrence. The coda Dθ\mathcal{D}_\theta is just the prediction head: normalise, modulate, project, unpatchify. It reads out a velocity from any state and does not feed anything back.

Two details carry weight. The conditioning is c=et(t)+ey(y)c = e_t(t) + e_y(y) (time plus class) everywhere, but the core alone also gets es(sk)e_s(s_k), an embedding of a continuous depth coordinate, through the same adaptive-norm and gate machinery DiT uses for time. And the coda has zero transformer blocks. The reason is accounting. Training reads out every state, so a coda with LDL_D blocks would run K+1K+1 times per forward pass. The paper's Eq. 30 puts the extra training cost at K(LDfblk+fhead+fs)K(L_D f_{\mathrm{blk}} + f_{\mathrm{head}} + f_s). With LD=0L_D = 0, what's left is a few prediction heads and depth embeddings. Their Table 17 puts it at 0.0188% of the dense L/2 training budget.

Parameter counts follow the unique blocks. Dense DiT-L/2 is 24 blocks and 457.83M parameters. LiFT-L/2-R10 is 4 prelude blocks plus a 10-block core, 14 unique blocks and 270.24M parameters. Trained with Ktrain=2K_{\mathrm{train}} = 2, it executes 4+2×10=244 + 2 \times 10 = 24 blocks per forward pass, exactly the dense model's depth. The paper uses this to match training compute within a scale: same executed depth, same 500,000 updates at batch 256, fewer unique weights.

The target, which is the whole paper

The obvious way to train this is to ask every readout u1,…,uKu_1, \dots, u_K to predict u⋆u^\star. This is deep supervision, what ELT and the Deep Supervision in Looped-DiT do, and its failure is easy to see. If loop 2 is already supposed to output the final answer, loop 3 has no defined job. Training never shows it one. Past the training depth the model is either at a fixed point (extra loops do nothing) or drifting off one (extra loops hurt). The related-work numbers the paper cites fit that: FID worsening as soon as inference exceeds training depth on ImageNet, quality peaking at 1.5 times training depth on video.

LiFT gives each loop a different target. Take the readout of the prelude state, u0u_0, stop its gradient, and call it the anchor b=sg(u0)b = \mathrm{sg}(u_0). Then draw a straight line from the anchor to the flow-matching target:

uˉs=(1−s) b+s u⋆,s∈[0,1].\bar u_s = (1-s)\,b + s\,u^\star, \qquad s \in [0,1].

Loop kk runs at depth coordinate sks_k and is regressed onto uˉsk\bar u_{s_k}:

LLiFT(θ)=Ex0,x1,t,S[1Ktrain d∑k=1Ktrain∥uk−((1−sk) b+sk u⋆)∥22].\mathcal{L}_{\mathrm{LiFT}}(\theta) = \mathbb{E}_{x_0, x_1, t, \mathcal{S}}\left[\frac{1}{K_{\mathrm{train}}\, d}\sum_{k=1}^{K_{\mathrm{train}}} \big\lVert u_k - \big((1-s_k)\,b + s_k\,u^\star\big)\big\rVert_2^2\right].

Put this next to the first equation in the article. xt=(1−t)x0+t x1x_t = (1-t)x_0 + t\,x_1 is a straight line from a starting point to a target, indexed by a continuous coordinate. uˉs=(1−s)b+s u⋆\bar u_s = (1-s)b + s\,u^\star is a straight line from a starting point to a target, indexed by a continuous coordinate. The time axis runs from noise to an image. The depth axis runs from the model's first guess at the velocity to the true velocity. TT sets how finely you walk the first line, and KK sets how finely you walk the second. This is the correspondence in Derakhshani's post, and the reason the paper can say both counts "can be chosen after training".

Each loop's job is now a fraction of a correction. Subtract the target from the reference and you get

u⋆−uˉs=(1−s) (u⋆−b),u^\star - \bar u_s = (1-s)\,(u^\star - b),

so a loop at s=0.25s = 0.25 is asked to remove a quarter of the gap between the first guess and the answer, and leave the rest. The last loop always sits at s=1s = 1, so the final readout is trained on the ordinary flow-matching loss. Nothing about the sampler changes.

reference path · LiFT Eq. 4-6
b = sg(u₀), s = 0u* = x₁ − x₀, s = 1k=1k=2k=3s
loop 1 target
s = 0.389
left for loop 1 to fix
61% of u* − b
average share per loop
1/3 = 33.3% of the correction

Training: K_train − 1 interior coordinates are drawn uniformly and sorted; the last is always s = 1, the plain flow-matching target. Each draw gives the same loops a different set of jobs.

Switch the widget to "every loop → u*" to see deep supervision: all the targets pile onto one point, and extra loops have nowhere to go.

Does the line make sense as a regression target, given that u⋆u^\star is noisy and the model can only learn its conditional mean? The paper's Appendix C.1 answers that in two lines. Write z=(xt,t)z = (x_t, t), hold the anchor function fixed, and take the conditional expectation:

E[uˉs∣z]=(1−s) b(z)+s v⋆(z),\mathbb{E}[\bar u_s \mid z] = (1-s)\,b(z) + s\,v^\star(z),

E[∥f(z)−uˉs∥2∣z]=∥f(z)−E[uˉs∣z]∥2+s2 E[∥u⋆−v⋆(z)∥2∣z].\mathbb{E}\big[\lVert f(z) - \bar u_s\rVert^2 \mid z\big] = \big\lVert f(z) - \mathbb{E}[\bar u_s \mid z]\big\rVert^2 + s^2\,\mathbb{E}\big[\lVert u^\star - v^\star(z)\rVert^2 \mid z\big].

The optimum at each depth slides linearly from the first guess to the marginal velocity, and reaches the ordinary flow-matching optimum at s=1s = 1. The second term is noise the network can't remove, and it scales with s2s^2. That term comes back later.

Why a straight line, and not some learned curve? Appendix C.3 shows the straight path is the unique minimiser of the kinetic action 12∫01∥γ˙(s)∥2 ds\tfrac12\int_0^1 \lVert\dot\gamma(s)\rVert^2\,\mathrm{d}s between the two endpoints. The same holds for the piecewise-linear version on any grid, so a randomised grid moves the knots but not the path. The Gaussian, Fisher and KL readings in C.4 all land on the same MSE. I read those as reassurance rather than motivation. The practical argument is simpler: a straight line is the only target whose knots you can place anywhere in [0,1][0,1] without the path changing under you.

Two small things that hold it together

The anchor is detached on the target side only. Gradients still flow through every prediction, through the whole rollout and into the prelude. This keeps the target from chasing the model within a step, and it creates a failure mode. Nothing in LLiFT\mathcal{L}_{\mathrm{LiFT}} asks u0u_0 to be a good velocity, so its scale can drift. If bb grows, the (1−s) b(1-s)\,b term dominates every intermediate target and the core has to learn an enormous correction.

It does drift. In the B/2 ablation with no extra term, the anchor's RMS climbs to 3.39 against a target RMS of 1.30, and its MSE to the target reaches 7.96. The fix is a small auxiliary loss, λ E∥u0−u⋆∥2/d\lambda\,\mathbb{E}\lVert u_0 - u^\star\rVert^2/d, with λ=0.01\lambda = 0.01, applied once outside the loop average. That holds the anchor RMS near 0.93. Bigger weights help the prelude's own error and hurt generation: FID 64.81 at λ=0.01\lambda = 0.01 against 70.03 at λ=1\lambda = 1 on that small model. So the prelude is nudged toward being a decent first guess and no further.

Why the loops extrapolate

The second ingredient is that ss is continuous and random during training. For each example at each update, LiFT draws Ktrain−1K_{\mathrm{train}} - 1 coordinates from U(0,1)\mathcal{U}(0,1), sorts them, and appends s=1s = 1. A model trained with Ktrain=2K_{\mathrm{train}} = 2 therefore runs its core twice per example: once at a random interior point, once at the endpoint. At inference the grid is uniform, sk=k/Kinfs_k = k/K_{\mathrm{inf}}.

The thread says "the model does not learn 'loop 1,' 'loop 2,' or 'loop 3.' It learns where it is along a continuous correction trajectory." True, and it undersells the result. At 16 inference loops, every ss the model sees is one it trained on. What it has never seen is a history 16 applications long. Training always had exactly two. The coordinates interpolate; the recurrence extrapolates.

The ablation in Appendix B.1 isolates the random draw. Two B/2 models trained with K=4K = 4, one on random coordinates and one on the fixed grid sk=k/4s_k = k/4, everything else equal and λ=0\lambda = 0:

KinfK_{\mathrm{inf}}12481632
random ss (FID)98.4878.5768.5166.9767.7168.38
fixed ss (FID)151.2399.4768.3166.6075.2686.95

At the training depth the two are identical. The fixed grid edges ahead at 5 and 8 loops, then falls apart, 86.95 at 32. The random grid stays flat. It is also far better below training depth, which matters for the use case the authors pitch, running shallow when compute is short.

Appendix B.3 checks whether the loops actually do what they were told. Take L/2 R10 (trained with two loops), project each intermediate readout onto the direction from bb to u⋆u^\star, and measure how far along it got. At four loops the mean projections are 0.246, 0.493, 0.737 and 0.982 at s=0.25,0.5,0.75,1s = 0.25, 0.5, 0.75, 1. At 32 loops, at the same coordinates: 0.249, 0.495, 0.738, 0.982. The model walks the line it was trained on, at sixteen times its training depth, with the off-line residual growing only from 6.07% to 8.68% of the prediction norm. I found this more convincing than any FID table in the paper, because it shows the loops doing the job the target defines.

Three panels, B/2, L/2 and XL/2, plotting FID on a log scale against inference loops K_inf from 1 to 32 at 50 steps. Each panel has several LiFT curves, one per core size, with rings marking training depth, and dotted horizontal lines for dense B/2, L/2 and XL/2. In L/2, the R10 curve trained at K=2 drops from about 30 to about 11 by 8 loops, well below dense XL/2 near 15.5. In XL/2, R12 and R8 fall below 10. In B/2, every LiFT curve stays above dense B/2 near 30. Single-block cores (R1) flatten near their training depth.
FID against inference loops at 50 steps. The large-core L/2 and XL/2 checkpoints keep improving well past their training depth (rings) and cross below the dense models; single-block cores flatten, and nothing at B/2 reaches dense B/2 (LiFT paper, Figure 3).

The gains are not uniform, and the paper is upfront about it:

Checking "52% fewer inference FLOPs"

The headline in the abstract and the thread: LiFT-L/2 beats the dense DiT-XL/2 baseline by 3.34 FID with about 60% fewer parameters, 32% fewer training FLOPs and 52% fewer inference FLOPs. I checked every ratio against Tables 3 and 15. The settings are LiFT L/2 R10 at 50 steps and 8 loops (FID 10.95, 28.24 TFLOPs per image, 270.24M parameters, 61.98 EFLOPs to train) against dense XL/2 at 250 steps (FID 14.30, 59.31 TFLOPs, 674.82M, 91.10 EFLOPs). Then 1 − 270.24/674.82 is 60%, 1 − 28.241/59.308 is 52%, and 1 − 61.984/91.098 is 32%. All three are correct.

What the headline leaves out is the operating point it picked. Dense XL/2 at 250 steps is on its plateau. The same model at 100 steps gets 14.69 for 23.72 TFLOPs, and at that point LiFT's 28.24 TFLOPs is about 19% more compute, not 52% less. LiFT still wins on quality by a wide margin (10.95 against 14.69), with 60% fewer parameters and 32% less training. So the quality and parameter claims stand. The 52% comes from comparing against a dense model spending most of its compute on steps that no longer help, which is the paper's own point about steps saturating, used to make the saving look bigger.

The comparisons I'd lead with are less dramatic and fairer:

Three panels, B/2, L/2 and XL/2, plotting FID on a log scale against inference TFLOPs per image on a log axis. Dotted dense curves vary integration steps from 1 to 250; solid LiFT curves vary loops at 50 steps, with rings at training depth. In L/2 and XL/2 the best LiFT curves continue below the dense curves' floor at similar or lower cost. In B/2 the dense curve stays below every LiFT curve.
Looping more against sampling more, on cost. At L/2 and XL/2 the large-core LiFT curves go below where the dense models flatten out; at B/2 the dense model is better at every cost (LiFT paper, Figure 4).

There are two more things the headline doesn't show. At its training depth, every LiFT checkpoint is worse than the same-scale dense model at the same cost: L/2 R10 scores 20.30 at two loops, dense L/2 scores 16.94. The paper explains this with an iso-depth result from looped language models, where a shared block is worth about half a unique one. The win only appears once you spend extra inference compute on loops. And at B/2 it never appears. The best LiFT B/2 result, 33.87, doesn't reach the 32.50 that dense B/2 gets at 25 steps. The authors' explanation is that the B/2 cores, at most 89M parameters, are too small for 32.8B training tokens. Plausible, and untested.

How to split a budget between steps and loops

So loops beat steps sometimes. When? The paper sweeps L/2 R10 over T∈{10,25,50,100}T \in \{10, 25, 50, 100\} and Kinf∈{1,2,4,8,16,32}K_{\mathrm{inf}} \in \{1, 2, 4, 8, 16, 32\}, 24 settings at 50,000 samples each. That grid is the most useful thing in the paper if you would actually deploy one of these.

budget split · LiFT L/2 R10, Tables 3 and 6FID, lower is better · 270M params
T \ K12481632
10
48.6
0.94 TF
33.6
1.61 TF
19.0
2.96 TF
15.3
5.65 TF
15.1
11 TF
16.4
22 TF
25
34.8
2.35 TF
22.6
4.04 TF
13.2
7.40 TF
11.4
14 TF
11.4
28 TF
12.1
54 TF
50
31.8
4.71 TF
20.3
8.07 TF
12.3
15 TF
11.0
28 TF
11.0
55 TF
11.6
109 TF
100
30.6
9.41 TF
19.4
16 TF
12.0
30 TF
11.0
56 TF
11.0
110 TF
11.5
218 TF
best LiFT setting that fits
T=25, K=4 → FID 13.16 (7.397 TF)
LiFT with 1 loop, steps only
T=50 → FID 31.83
dense L/2 (458M), steps only
T=50 → FID 16.94 (8.069 TF)
dense XL/2 (675M), steps only
T=25 → FID 17.68 (5.931 TF)

Faded cells cost more than the budget. All settings use Euler sampling with no classifier-free guidance; dense L/2 trains with the same FLOPs as this checkpoint, dense XL/2 with about 47% more. Rows are integration steps T, columns are loops per step K.

Two rows of that table make the case for compute inside each step better than any prose. Hold loops at one and add steps: 48.61 at 10 steps, 30.59 at 100. Hold steps at ten and add loops: 48.61 at one loop, 15.29 at eight. The sampler can't fix a bad velocity estimate however many steps you give it. The loops fix the estimate.

It also cuts the other way. Sixteen loops at ten steps gets 15.09 for 11.027 TFLOPs, which is both worse and more expensive than four loops at 25 steps (13.16 at 7.397). Past a few loops, a ten-step Euler sampler is the bottleneck again, and the next unit of compute should go to steps. At the top end both axes flatten: eight loops gives 10.95 at 50 steps and 10.96 at 100.

Two panels. (a) Best measured FID within each inference budget, step plots against TFLOPs per image on a log axis: LiFT L/2 R10 in orange descends to about 11, while dense L/2 levels off near 16 and dense B/2 near 29; annotations mark (25, 4) and (50, 8). (b) A heatmap of FID over recurrent depth K_inf (1 to 32, x axis) and integration steps T (10 to 100, y axis), with contour lines, dashed lines of equal cost, a white line marking the cost of dense L/2 at 50 steps, a circle at the training setting (2 loops, 50 steps), an orange star at 4 loops and 25 steps labelled 13.16, a black star at 8 loops and 50 steps labelled 10.95, and a pink path tracing the best allocation as budget grows.
Left: the lowest measured FID inside each budget. Right: FID over loops and steps; the star at (25 steps, 4 loops) sits below dense L/2 on the cheap side of its cost line, and the best-allocation path moves along both axes, never just one (LiFT paper, Figure 5).

The pink path in panel (b) is the real recipe. As the budget grows it adds loops at ten steps, trades some of them for steps, adds loops again, and finally adds steps. If I were running this I'd pick a small set of (T, K) presets off that path and expose them as one quality knob. A single checkpoint covers the whole range, which is the property the authors care about most.

Velocity MSE barely moves; FID does

There's a number in Appendix B.4 that I think matters more than the paper suggests. Take L/2 R10, feed the same noisy inputs through grids of different depth, and measure the MSE of the final velocity against u⋆u^\star. For every grid with K≥2K \ge 2 it sits between 0.7667 and 0.7689. FID over the same range of depths goes from 20.30 to 10.95.

No contradiction: it's the s2s^2 term from the decomposition above. At s=1s = 1, most of the per-sample error against u⋆u^\star is the conditional variance of u⋆u^\star given zz, noise no network can remove, and it swamps the part the loops improve. The practical consequence: if you train something like this, validation velocity MSE will tell you almost nothing about whether more loops help. You have to sample. The paper does, with 50,000 images per setting.

What the pictures show, and what they don't

A grid of ImageNet samples, seven rows (rooster, macaw, golden retriever, peacock, tabby cat, zebra, red panda) and six columns. The left three columns are LiFT L/2 R10 at K_inf 2 (training depth, 8.07 TFLOPs), 8 (28.24) and 16 (55.14), all at 50 steps; the right three are dense L/2 at 50, 100 and 250 steps (8.07, 16.14, 40.35 TFLOPs). Each row shares class and noise. At 2 loops the retriever and peacock are malformed; at 8 and 16 loops the retriever is a coherent dog and the peacock fans its tail. The dense columns change little between 50 and 250 steps; their retriever and peacock stay malformed.
Paired samples from the same noise: LiFT L/2 R10 at 2, 8 and 16 loops against dense L/2 at 50, 100 and 250 steps. The paper calls these selected samples (LiFT paper, Figure 7).

Look at the retriever and peacock rows. At its training depth LiFT produces the same kind of mangled animal the dense model does. At eight loops it is a dog, while more steps barely change the dense columns. The paper's whole argument fits in those two rows. But the caption says selected, so read it as an illustration and take the 50,000-sample FIDs as the evidence.

I tried to make my own strip through the demo with a different class, same seed, 1 to 16 loops. The first call came back, and then the Space's ZeroGPU quota for anonymous callers ran out for the day. So I have nothing to add here beyond the authors' images. The demo exposes class, seed, integration steps and KinfK_{\mathrm{inf}} (1 to 32) for the L/2 R10 checkpoint, plus dense L/2 and XL/2 at up to 250 steps. Its footer says "500k-step EMA checkpoints · SD-VAE · Euler sampling · no CFG", which matches the paper's protocol.

Where it sits among the other looped DiTs

Three neighbours make LiFT's choice clearer.

So the architecture isn't new. Prelude, shared core, depth embedding and input reinjection are all borrowed, and the paper says so. What's new is giving each loop a different target, a continuous coordinate for it, and the observation that this reuses flow matching's own construction. The authors call it "only light changes to the standard architecture", which is accurate.

What I'd want before using it

The results are on class-conditional ImageNet 256×256 with no classifier-free guidance. Guidance changes the velocity field a lot, and whether the depth curve survives it is the first open question. The authors list it as a limitation. The absolute FIDs are unguided numbers from a 500,000-update budget. They are fine for comparing these models with each other and say nothing about state of the art.

Every number comes from one training run and one sample set. The FID gaps that matter here are several points, much larger than seed noise usually is, so I'm not worried about the main claims. I would be worried about the 0.04-point differences between neighbouring settings.

Costs are analytic FLOPs. There is no latency, memory or energy measurement, and the paper says so. For this architecture FLOPs should track wall clock better than usual: every loop is a full-width block with no KV cache to carry, so a loop costs what a block costs. But the edge-device pitch in the conclusion needs a device to back it.

And the code isn't out. The tweet promises "the codebase, W&B artifacts, and checkpoints"; the paper's links say "Coming Soon". The method is simple enough that Algorithm 1 is close to a spec. Here it is in PyTorch, transcribed from the paper's pseudocode (Algorithm 1, lines 1 to 14). The prelude, core, head and embedding modules are placeholders:

# One LiFT training step, transcribed from Algorithm 1 of arXiv 2610.05538.
def lift_loss(x1, y, K, lam=0.01):
    B = x1.shape[0]
    x0 = torch.randn_like(x1)
    t = torch.rand(B, device=x1.device)
    xt = (1 - t.view(B, 1, 1, 1)) * x0 + t.view(B, 1, 1, 1) * x1
    u_star = x1 - x0
 
    # K-1 interior depths per example, sorted, then the endpoint s=1
    s = torch.sort(torch.rand(B, K - 1, device=x1.device), dim=1).values
    s = torch.cat([s, torch.ones(B, 1, device=x1.device)], dim=1)
 
    c = e_t(t) + e_y(y)
    h0 = prelude(xt, c)
    u0 = head(h0, c)
    b = u0.detach()                                   # anchor: stop-gradient
    loss = lam * F.mse_loss(u0, u_star)               # prelude loss, outside the average
 
    h = h0
    for k in range(K):
        sk = s[:, k]
        h = core(rms_norm(h) + h0, c + e_s(sk))       # reinject h0 every pass
        uk = head(h, c)
        target = (1 - sk.view(B, 1, 1, 1)) * b + sk.view(B, 1, 1, 1) * u_star
        loss = loss + F.mse_loss(uk, target) / K
    return loss

Sampling is the same loop with sk=k/Kinfs_k = k/K_{\mathrm{inf}}, no intermediate readouts and one Euler update per step. The paper also nudges the sorted draws apart by 10−610^{-6} in FP64 so no two coincide; I left that out.

Would I use it? For a deployment where weights are the constraint and you want one checkpoint covering several quality and latency budgets, yes. 270M parameters doing better than a 675M model is a real saving in memory. For a server that only ever runs one budget, the case is weaker. A dense model at its best step count is a known quantity, and LiFT's advantage there is the cleaner 8 to 38% compute figures above, not 52%.

The idea I expect to outlast this paper is the target. "Give each loop a fraction of the correction, indexed continuously" doesn't depend on images or on flow matching's endpoint. Any looped model with a prediction at every pass, language models included, could try the same thing. The open question is whether there is an anchor as natural as sg(u0)\mathrm{sg}(u_0) when the output is a distribution over tokens.

How I checked

I read the arXiv HTML of 2610.05538v1 in full, including the appendices, and took every number in this piece from its tables (3, 4, 5, 6, 8, 9, 10, 11, 12, 15, 17) or its text. The X threads by both first authors were read through the fxtwitter mirror; they add the demo link and nothing that contradicts the paper.

I re-derived the FLOP accounting from the paper's Eq. 32, fblk=n(8w2+4wm+4nw)+12w2f_{\mathrm{blk}} = n(8w^2 + 4wm + 4nw) + 12w^2 with n=256n = 256 tokens and m=4wm = 4w. For L/2 (w=1024w = 1024) that gives 6.7235 GFLOPs per block per image. Dense L/2 at 50 steps comes out at 8.068 TFLOPs against the table's 8.069. LiFT L/2 R10 at 8 loops (4+80=844 + 80 = 84 blocks) comes out at 28.239 against 28.241. Dense XL/2 at 250 steps comes out at 59.30 against 59.308. Training compute 3⋅U⋅B⋅Cforward3 \cdot U \cdot B \cdot C_{\mathrm{forward}} gives 61.96 EFLOPs for dense L/2 against 61.973. The small differences are the heads and embeddings I left out. The headline ratios (60%, 52%, 32%) and my alternative comparisons are plain divisions of table entries. One rounding note: 14.30 − 10.952 is 3.348, which rounds to 3.35, not the paper's 3.34. That only matters if the dense value printed as 14.30 is slightly below it, and either way it's the same result.

The figures are the paper's own: Figure 2 and Figures 3 to 5 were SVGs, which I rasterised at 1600 px on white; Figure 7 is the paper's PNG. The two widgets use the paper's equations and Tables 3 and 6 directly, with no interpolation. I read the demo's Gradio config to see what it exposes, and one call succeeded before the anonymous GPU quota ran out. I couldn't check anything that depends on code, since none is released.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "LiFT: a looped DiT that keeps improving past the loops it was trained with", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026liftloopflowtransformer,
  author = {Satyajit Ghana},
  title  = {LiFT: a looped DiT that keeps improving past the loops it was trained with},
  url    = {https://ai.thesatyajit.com/articles/lift-loop-flow-transformer},
  year   = {2026}
}
share