~/satyajit

LOOM and Looped-DiT: looping a model more than twice

mdjsonmcp

2026-10-06 · 21 min · looped-transformers · recurrent-depth · mixture-of-experts · diffusion · image-generation · architecture · llm · explainer

When this site covered SMELT, the main finding was a ceiling. Under matched FLOPs, parameters and KV cache, looping a MoE transformer's middle layers twice beat the unlooped baseline. Looping it three or four times did worse. A concurrent paper, "Loop the Loopies!", reached the same optimum of two under wall-clock matching. Nanbeige4.2-3B shipped with a loop count of 2. Three groups, one answer.

Two papers from the last week push past it:

Looping is depth with shared weights

A standard transformer applies MM distinct blocks once each. A looped transformer applies the same MM blocks HH times:

ht=Fθ ⁣(ht−1),t=1,…,H,h0=Embed⁡(x)h^{t} = F_\theta\!\left(h^{t-1}\right),\quad t = 1,\dots,H,\qquad h^{0} = \operatorname{Embed}(x)

Effective depth is MHMH, and the parameter count is still that of MM blocks. That is the whole appeal. The looped transformer architecture page covers the lineage: Universal Transformer, ALBERT, Huginn and Ouro.

The catch is that a model has three budgets, and looping moves only one of them:

  1. Parameters stay flat. That is the point of sharing weights.
  2. FLOPs per token grow about HH-fold, because every pass is a full forward pass through the shared blocks.
  3. KV cache also grows about HH-fold in an autoregressive model. Each pass runs attention on a different hidden state, so each pass writes its own keys and values.

A parameter-matched comparison is therefore a compute-mismatched one, which was this site's complaint about virtual logic depth. Huginn versus Ouro shows how many other design axes hide inside the word "looped". The first question for any looped result is which budget was held fixed.

Why two loops is where naive looping stalls

LOOM's diagnosis has two parts.

Variance growth (the curse of depth). In a pre-norm residual stream, hℓ=hℓ−1+Fℓ(Norm⁡(hℓ−1))h_{\ell} = h_{\ell-1} + F_\ell(\operatorname{Norm}(h_{\ell-1})). The norm keeps every update roughly the same size while the stream grows as updates pile up, so each new update moves the state relatively less and deep layers drift toward identity maps. Sun et al. called this "the curse of depth" (Shiwei Liu is an author on both papers).

Looping makes it worse, because the updates are no longer independent. Ten different layers write roughly uncorrelated vectors, so their sum grows like 10\sqrt{10}. The same layer writing on its ninth visit tends to push in the same direction it pushed on visits one through eight, so the sum grows closer to linearly in HH. LOOM's Figure 3 shows this directly. On a 9-loop, ~350M model with no fixes, activation variance sits in the hundreds. With residual scaling it stays near 1 (reported).

Expert-selection collapse. In a looped MoE, the router for layer ℓ\ell is part of the shared block. The hidden state changes only a little between loops, so the router sends the token to nearly the same experts each time. The extra loop then re-runs a similar computation instead of a new one. LOOM's Figure 4 measures this as the cosine similarity of per-loop expert-load distributions. With one shared router, every pair of loops sits near the top of the scale, which runs from 0.88 to 1.00.

Without the variance fix, extra loops destabilize training. Without the diversity fix, they are stable but redundant. Either way the third loop does not earn its FLOPs.

The LOOM recipe

LOOM's stated principle is "each loop should contribute new computation while keeping the recurrent state stable." Three components keep the state stable and two make the loops differ:

LOOM architecture diagram. Three columns show loop 1, loop 2 and loop H of an M-layer stack. Each layer is a dashed box holding 'Looped MoE L_m' above an orange 'Router R^t_m' whose superscript is the loop index, so every loop has its own routers. EMA boxes sit between layers. Curved arrows carry the top of each loop into an EMA box at the bottom of the next loop. A line from the Embedding block at the bottom left feeds each loop through 'Scaling g_1' to 'Scaling g_H'. On the right, one layer is expanded: Norm, Attention, Scaling gamma, EMA, Norm, MoE, Norm, Scaling gamma.
LOOM's architecture: an M-layer MoE block reused for H loops. Every loop has its own routers (orange), an EMA Looping Residual carries attention outputs forward, the input embedding is re-injected at each loop with weight g_t, and both branches are scaled by γ (LOOM paper, Figure 2).

Residual scaling. Both the attention and MoE branch outputs are multiplied by

γ=λHM,λ=0.5\gamma = \frac{\lambda}{H\sqrt{M}}, \qquad \lambda = 0.5

The two factors match the two kinds of accumulation. Across the MM layers of one loop, updates are roughly independent, so their sum grows like M\sqrt{M}, and 1/M1/\sqrt{M} cancels that. Across the HH loops, the same layer's updates are correlated, so their sum grows like HH, and 1/H1/H cancels that. Multiplied out, the total stays near λ\lambda times one update's scale, whatever the loop count (reasoned, from the paper's Eq. 1). The factor itself is from Wang et al.; LOOM applies it to MoE.

Embedding re-injection. At the start of each loop t≥2t \geq 2, the state is mixed back toward the input embedding:

h0t=(1−gt) hMt−1+gt x,gt=λtMh_0^{t} = (1-g_t)\,h_M^{t-1} + g_t\,x, \qquad g_t = \frac{\lambda}{t\sqrt{M}}

For the 10-layer model, g2≈0.079g_2 \approx 0.079 and g9≈0.018g_9 \approx 0.018 (reasoned). The anchor is strongest early and fades, so late loops mostly refine their own state. Huginn re-injects the input too; the decay is new.

Loop-specific routers. Each loop gets its own router at every layer, while the expert weights stay shared. Router RℓtR^t_\ell can send the same hidden state to different experts on loop 3 than on loop 7. In Figure 4 the shared-router matrix is nearly flat near the top of the scale. The per-loop matrix is visibly checkered, with loop pairs dropping to the bottom of its 0.88-1.00 colour range. The README's weave animation puts a number on it. Over 9 loops, one token of the 1.7B model touches 14 distinct experts at one layer, against 6 if every loop reused the same ones. Averaged over 16,384 held-out tokens it is 14.6, and two consecutive loops share about 3 of their 6 routed experts (reported, README).

Two 9-by-9 heatmaps of cosine similarity between loops' expert-load distributions, colour scale from 0.88 to 1.00. The left 'Independent' map, loop-specific routers, is checkered, with many off-diagonal cells in darker orange near 0.90. The right 'Shared' map is almost uniformly pale, near 0.97 to 1.00 everywhere.
Cross-loop similarity of expert-load distributions in a 9-loop ~350M model, loop-specific routers (left) versus one shared router (right); lower means more diverse expert use (LOOM paper, Figure 4a).

The cost is small: at d=512d=512, E=80E=80, M=10M=10, five loops carry 5×10×512×80≈2.05 \times 10 \times 512 \times 80 \approx 2.0M router parameters in a 700M model (reasoned).

Looping Residual. This is the part that is new to LOOM. The direct attention residual is replaced by two exponential moving averages of the scaled attention outputs, each kept as an accumulator plus a scalar normalizer:

N←βN+oℓt,D←βD+1,r=N/D,β=0.5N \leftarrow \beta N + o_\ell^t,\quad D \leftarrow \beta D + 1,\quad r = N/D,\qquad \beta = 0.5

The global memory persists across all layers and loops. The local one resets at the start of each loop. The update becomes h~=h+rH+rL\tilde h = h + r_H + r_L. With β=0.5\beta = 0.5, the normalizer tends to 2, so the newest attention output carries about half the weight, the one before it a quarter, and so on (reasoned). Later loops see a summary of earlier loops' attention for one tensor and one scalar per memory. The paper writes both memories as fed by the raw attention output. The shipped scripts/train.sh sets arch.dual_axis_h_update_src=l, which feeds the global memory from the local summary instead (measured, models/dual_axis_carry.py line 823). Small, but the equation and the code that produced the numbers differ.

MoE output RMSNorm. The MoE branch output is RMS-normalized before γ\gamma scales it. The paper treats it as plumbing; the ablation says otherwise.

Three line plots for a 9-loop ~350M MoE. Left: training loss over 10,000 steps. LOOM falls smoothly to about 2.7. Only Residual Scale falls to about 3.4 with a spike near step 7,000. Native and Only Embedding Inject plateau near 7. Middle: activation variance during training on a log scale. Native and Only Embedding Inject climb into the hundreds to thousands, while LOOM and Only Residual Scale stay near 1. Right: activation variance across loop iterations 1 to 9. Embed-only sits above 1,000, Native near 400, LOOM and Residual Scale near 1.
Residual scaling controls variance but alone still trains to a worse loss; the full recipe gets both (LOOM paper, Figure 3: ~350M model, 9 loops, 10B tokens).

Segmented backprop, and the optimizer steps it adds

Backpropagating through 12 loops of an MoE stores 12 loops of activations. LOOM splits the HH loops into segments of at most K=3K=3. Each segment ends in an LM-head loss and a backward pass, and the hidden state and global memory are passed forward to the next segment with gradients detached. On the ~350M model with micro-batch 1, peak memory stays near 12.6 GiB from 3 to 12 loops, against 12.5 to 19.3 GiB without segmentation, and 12-loop training time drops from 26.1 to 12.6 hours (reported). Unsegmented 9-loop training also diverges: Table 6 lists a perplexity of 81.04 with an unrecoverable spike.

Algorithm 1 has one detail the prose does not spell out: line 10 reads "backward; optimizer step", and it runs inside the loop over tt. I checked the code. run_segmented_global_step in pretrain.py calls step_optimizers once per segment (measured). So a 9-loop model takes 3 optimizer updates per global batch and a 12-loop model takes 4, while the non-looped baseline takes 1. All of them see the same tokens and run the same learning-rate schedule, which is keyed to global steps. Each segment's loss also supervises loops 3, 6 and 9, not only the last. So segmentation is a memory trick and an optimizer change. The non-iso-FLOP results use it; the iso-FLOP results do not.

The numbers, with the arithmetic

loop ledger · LOOM tables 2-4

M=10, 80 routed experts, 10B tokens. top-k shrinks as loops grow so f(H,k) stays near 84. Full backprop, no segmentation.

effective depthpaper
10 x 5 = 50
per-token computereasoned
f = 90 d² (top-4)
KV cachereasoned
5x baseline
optimizer steps / batchcode
1 (full backprop)
validation PPLreported
16.54 (-1.82)
7-task avg accreported
39.53% (+0.69)
attention 33%experts 67%

Where the fixed budget goes (reasoned from Eq. 5): attention costs 6 per pass, experts 3 per selected expert per pass. Looping at iso-FLOP moves compute out of experts and into attention.

18.361x17.372x16.913x16.584x16.545x16.576x

Filled: LOOM validation perplexity, lower is better.

Iso-FLOP: best at 5 loops, and why that is a fair claim

The 700M experiment (Table 2) holds parameters fixed: M=10M=10, d=512d=512, 80 routed experts of width 512. It pays for extra loops by cutting top-kk. LOOM counts multiply-adds per token per layer in units of d2d^2. Attention is 6d26d^2: 33 for grouped-query QKV, 22 for causal SDPA and 11 for the output projection. Each SwiGLU expert is 3d23d^2:

f(H,k)=H (6+3k)f(H,k) = H\,(6 + 3k)

I checked every row: 1×(6+3⋅26)=841 \times (6 + 3 \cdot 26) = 84, 2×(6+36)=842 \times (6 + 36) = 84, 3×(6+24)=903 \times (6 + 24) = 90, 4×(6+15)=844 \times (6 + 15) = 84, 5×(6+12)=905 \times (6 + 12) = 90, 6×(6+9)=906 \times (6 + 9) = 90. They all match the table (reasoned). The SDPA term checks out too: at sequence length 1,024 =2d= 2d, a causal token attends to 512=d512 = d keys on average, so QK⊤QK^\top and AVAV cost d2d^2 each. The uncounted router, about 0.16 d20.16\,d^2 per loop, is under 1% of the budget (reasoned).

LoopsEff. depthtop-kf(H,k)Val. PPL7-task avg
1 (baseline)10268418.3638.84%
220128417.3739.00%
33089016.9139.34%
44058416.5839.53%
55049016.5439.53%
66039016.5739.50%

Reported, LOOM Table 2. 700M parameters, 10B tokens, full backprop.

The headline row, 5 loops, spends 90/8490/84, about 7% more FLOPs than the baseline (reasoned). The 4-loop row is exactly matched at 84 and reaches 16.58, within 0.04 of it. So "at matched FLOPs, looping beats the baseline, and the optimum is past two" survives the stricter reading: 18.36 → 16.58 at exactly equal cost (reported numbers, reasoned comparison). These runs use full backprop, so the segmentation confound does not apply. On the downstream average the gain is small: +0.69 points at 4 or 5 loops, and the paper reports no seed variance, so one run per setting cannot establish it. The perplexity gain is what this table really shows.

Two dual-axis line plots. Left, 'Iso-FLOP, 700M (10B tokens)': validation PPL falls from about 18.4 at 1 loop to about 16.5 at 4 to 6 loops, while downstream average rises from 38.8 to about 39.5 and flattens. Right, 'Non-Iso-FLOP, 1.7B (60B tokens)': PPL falls from 9.6 at 1 loop to 7.8 at 9 loops and ticks up slightly at 12, while downstream average rises from about 42.4 to 47.7 at 9 loops.
Loop-depth scaling in LOOM: under near-iso-FLOP the 700M model bottoms out near 5 loops; without FLOP matching the 1.7B model peaks at 9 (LOOM paper, Figure 1).

The widget shows what the paper does not: the compute is moved, not added. In the baseline, attention takes 6/84≈7%6/84 \approx 7\% of the budget; at 5 loops it takes 30/90≈33%30/90 \approx 33\% (reasoned). Iso-FLOP looping swaps a wide expert mixture (26 of 80 experts per token, a very dense MoE) for a narrow one run five times, with attention every time. Some of the win could come from spending more on attention, and the paper has no non-looped control with that split. The KV cache also grows fivefold, from 10 attention passes per token to 50. LOOM reports no inference memory or decode latency, so this is a training-compute claim. SMELT matched the cache; LOOM does not.

Iso-param: best at 9 loops, at about 9x the cost

At 1.7B parameters (0.63B active) and about 60B tokens, Table 4 reports:

LoopsEff. depthVal. PPL7-task avg
1159.6242.4%
3458.9443.9%
6907.9146.7%
91357.7747.7%
121807.8447.5%

Reported, LOOM Table 4.

This is a real stability result: a 1.7B MoE trains through 135 effective layers and keeps improving up to 9 loops. On the ~350M backbone (Table 3), looping with no technique reaches a perplexity of 23.75 at 3 loops, already worse than the 20.07 baseline, and hits unrecoverable spikes with perplexity above 1,300 at 6, 9 and 12. LOOM's 9-loop run on that backbone reaches 14.80.

It is not a free lunch, and the paper's "unrolled scale of roughly 15B parameters (1.7B × 9)" invites a misreading. The 9-loop model costs roughly 9x the baseline's per-token compute in the looped blocks, to train and to serve, holds 9x the KV cache, and took 3 optimizer steps per batch to the baseline's 1 (reasoned; the optimizer count is measured from the code). The fair baseline is a model trained with 9x the compute, and the paper does not run one. SMELT's matched-compute frontier put the real saving from looping at 6.8-18.0% of training FLOPs. Read the 9.62 → 7.77 drop as "LOOM makes deep loops trainable", not as "loops are worth 9x".

Ablations

At ~350M and 9 loops, with every variant evaluated at step 5,000 (Table 5, reported):

VariantVal. PPLAvg
LOOM18.6239.8
w/o residual scaling24.6338.6
w/o embedding re-injection26.0539.2
w/o MoE output RMSNorm26.5039.0
w/o Looping Residual19.2539.3
w/o per-loop routers (shared router)19.8339.2

The stabilizers matter most: removing any one of them costs 6-8 perplexity points. The diversity components matter less: 0.63 points for the Looping Residual, 1.21 for the per-loop routers. Expert-selection collapse is real, but at step 5,000 it is the smaller problem. Whether that gap widens by the end of training is not reported.

Looped-DiT: the same idea inside a denoising step

A diffusion transformer already iterates over denoising steps. Looped-DiT adds a second loop inside each step, starting from MiniT2I, a minimal pixel-space MMDiT with 17 blocks (see the diffusion transformer architecture page). Those blocks are split into 6 pre-loop, 5 looped and 6 post-loop blocks, and the middle 5 run N=4N=4 times. That gives 6+5×4+6=326 + 5 \times 4 + 6 = 32 block applications from 17 blocks of weights (reasoned, matches the paper's "effective depth 32").

Looped-DiT method diagram. Right: input x_t flows through Pre-loop (first A blocks), then a dashed brown box of shared-weight Looped middle B blocks repeated from the 1st to the Nth loop, then Post-loop (last C blocks) to the final prediction and flow loss l_N. Dotted training-only arrows send each intermediate loop's output through a shared post-loop copy to its own flow loss l_1 through l_(N-1), labelled Deep Supervision. Left: Self-Modulating Attention, applied inside looped blocks only. The standard Q, K, V projections and scaled dot-product attention produce o_(i,h), which is multiplied by G_(i,h), either a learned sigmoid gate (Gated Attention) or the projection I minus v-hat v-hat-transpose (Exclusive Self-Attention), before the heads are concatenated and projected.
Looped-DiT: shared middle blocks repeat N times inside each denoising step; Deep Supervision decodes every intermediate loop through the shared post-loop blocks during training; Self-Modulating Attention regulates the attention update inside the loop (Looped-DiT paper, Figure 2).

Naive looping has its own failure here, and the paper measures it. A ridge-regression probe that recovers each image token's 2D position from its hidden state drops from R2=0.865R^2 = 0.865 after loop 1 to 0.5620.562 after loop 8 (reported). Repeated attention updates overwrite where each patch is: the vision version of LOOM's drift. The two fixes:

The ablation stacks up cleanly on B/32 (Table 4, four-benchmark average, reported): no loop 55.2, naive loop 56.2, plus Deep Supervision 57.5, plus XSA alone 58.5, both 59.1. XSA on a non-looped model does nothing (−0.1). On the looped model it adds +1.6 (Table 5). That fits the claim that it damps repeated updates, not attention in general.

Checking "6.5x larger" and "4.9x"

Three-panel teaser. (a) Bar charts: on CoReBench, Looped-DiT B/16 scores 53.5 against InternVL-U 48.2, Uni-CoT 45.8, UniLIP-3B 45.0 and CogView4 45.0, a +5.3 lead. On PRISM, Looped-DiT scores 67.0 against InternVL-U 63.5, CogView4 60.9, Uni-CoT 58.8 and UniLIP-3B 58.5, +3.5. (b) Inference efficiency against InternVL-U: parameters 1.7B versus 0.26B (6.5x fewer), compute 732 versus 150 TFLOPs (4.9x less), latency 2.1 s versus 0.65 s (3.2x faster). (c) The prompt 'EMBRACE CHANGE, DREAM BIG AND CREATE A BRIGHTER TOMORROW' rendered at loops 1 to 4. Loops 1 to 3 misspell words and loop 4 is correct, while InternVL-U's render is garbled.
Looped-DiT's headline: a 0.26B model against the 1.7B InternVL-U on two reasoning-heavy benchmarks, the cost comparison behind '6.5x' and '4.9x', and text rendering corrected across loops (Looped-DiT paper, Figure 1).

The abstract says a 260M looped model "can surpass a model 6.5× larger … while requiring 4.9× lower inference compute." Both ratios come from comparing against InternVL-U. In Figure 1(b) the parameter bars read 1.7 and 0.26 (billion), and 1.7/0.26=6.541.7 / 0.26 = 6.54. The compute bars read 732 and 150 TFLOPs per image, and 732/150=4.88732 / 150 = 4.88. Latency is 2.1 s versus 0.65 s on one H100, 3.2×3.2\times (reasoned from the reported bars). The numbers are consistent. The X post that circulated the paper says "4.9x fewer parameters". That is wrong: 4.9x is the compute ratio, and the parameter ratio is 6.5x.

Two caveats. "Surpass" is on the six-benchmark average, 71.5 versus 69.0 (Table 1); Looped-DiT does not win GenEval (87.4, against 90.3 for UniLiP-3B). And the paper's training-data appendix describes the MiniT2I setup (CC12M, then BLIP3o-60K, DALL-E 3 and ShareGPT-4o-Image), while the shipped configs/b16_pretrain.yml also pretrains on FLUX-Reason-6M and configs/b16_finetune.yml mixes in Fine-T2I at weight 0.116, about half the samples (measured). The B/16 headline model saw more, and more reasoning-flavoured, data. I could not confirm what the MiniT2I-B/16 row (66.4) was trained on. The clean evidence is the B/32 work, where data is fixed.

Loops versus denoising steps

A diffusion model already has a knob for inference compute: more denoising steps. Figure 6 compares 4 loops at 25, 35 and 50 steps against 1 loop at 50, 70 and 100 steps, matched on per-image FLOPs.

Four small line charts for DPG, PRISM, CoRe and Spatial, each plotting score against TFLOPs per image at 14, 20 and 28. The blue Looped-DiT Loop x4 line (25, 35, 50 steps) sits highest in every panel. The brown Looped-DiT Loop x1 line (50, 70, 100 steps) is in the middle, and the grey non-looped baseline (50, 70, 100 steps) is lowest. Extra steps barely lift the brown and grey lines.
At matched inference FLOPs, four loops at fewer steps beat one loop at more steps, on the same checkpoint and on a separately trained non-looped model (Looped-DiT paper, Figure 6, B/32).

The FLOPs do match. Table 8 gives 146 GFLOPs per denoising forward pass with one loop and 267 with four, a ratio of 1.83, not 4, because the pre- and post-loop blocks run once either way. Then 25 steps × 267 = 6,675 GFLOPs and 50 steps × 146 = 7,300. If each step runs two passes for classifier-free guidance, these become about 13.4 and 14.6 TFLOPs, the "14" column of the plot (reasoned). The looped arm spends slightly less, and in every panel 4 loops at 25 steps beats 1 loop at 100 steps, which costs twice as much. My reading: past a few dozen Euler steps, more steps re-solve the same ODE more finely, while more loops give each step more depth to work out the layout, as in the sign that is misspelled at loops 1-3 and right at loop 4.

Diffusion also avoids LOOM's serving problem: a denoiser keeps no KV cache across steps, so looping costs FLOPs and latency, not cache memory (reasoned). Table 2 is honest about where the bill lands. Looped-DiT B/32 scores 59.1 against 58.1 for an untied 32-block model with the same Deep Supervision, but costs about 1.54× the training compute of the plain deeper and wider baselines, and the paper says so.

What carries across

The two papers solve the same problem in different domains, and their fixes rhyme:

FailureLOOM (MoE LM)Looped-DiT (T2I diffusion)
State drifts or blows upResidual scaling by 1/(HM)1/(H\sqrt{M}) plus MoE output normXSA or gate damping attention updates
Loop forgets its inputDecaying embedding re-injectionProbe shows positions eroding; XSA keeps R2R^2 highest
Loops repeat each otherPer-loop routers, Looping ResidualNot addressed: no experts to route
Only the last loop learnsSegment losses at loops 3, 6, 9 (a side effect)Deep Supervision on every loop, by design

Both papers show that loop count is a property of the training recipe, not of the architecture. "Two loops is optimal" was true of naive recipes; these move the optimum to 4-5 under near-matched FLOPs (LOOM) and to 4 trained loops with useful exits on either side (Looped-DiT). Neither shows that deep loops are cheap. LOOM's deepest wins are paid for in compute and cache; Looped-DiT's in training compute, at 0.26B scale.

The runs I'd want next: LOOM's iso-FLOP grid with the KV cache matched too; the 1.7B 9-loop row against a non-looped model given the same total compute and optimizer steps; and Looped-DiT at billions of parameters. Both codebases are public (Apache-2.0 and MIT), so all three are runnable.

hed-ucas/LOOM@a2ee5a7 · snapshot 2026-10-06
tracked files
53
license
Apache-2.0
branch
main
tests
none found
source
572.2 kB
commit date
2026-10-05
source by language
Python553.9 kB(33)Shell18.3 kB(5)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at a2ee5a7 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow

shallow clone: counts describe the pinned tree, not the history

OpenSenseNova/Looped-DiT@65a7705 · snapshot 2026-10-06
tracked files
36
license
MIT
branch
main
tests
1 file
source
132.7 kB
commit date
2026-10-05
source by language
Python132.7 kB(22)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at 65a7705 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

For the routing basics see the mixture-of-experts architecture page, and for where expert routing time goes in practice, Halo's breakdown.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "LOOM and Looped-DiT: looping a model more than twice", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026loomloopedmoe,
  author = {Satyajit Ghana},
  title  = {LOOM and Looped-DiT: looping a model more than twice},
  url    = {https://ai.thesatyajit.com/articles/loom-looped-moe},
  year   = {2026}
}
share