~/satyajit

Level-of-Token Diffusion: a DiT that is told where the detail goes

mdjsonmcp

2026-10-06 · 29 min · diffusion · diffusion-transformers · image-generation · video-generation · flow-matching · inference

Why read this

Solidtop 85%

How LoT turns a layout into rectangle tokens a pretrained DiT can denoise, why images gain less than the token cut and video more, and what eval layouts leak.

  • Original analysis
  • A new technique
  • Concrete numbers to act on

Image & video generationNothing to runResearch paper

How this was scored
Is it new?
2 of 3: A real new idea, method or capability
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
0 of 3: Closed, nothing to run
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
1 of 3: Relevant for months
Does it affect many?
1 of 3: A specialist community
Only here?
2 of 3: A teardown or measurement few others did

Score 55 of 100, ranked 293 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Gordon Wetzstein opened his thread with a line I have been thinking about for a while: "Diffusion models spend the same compute on a blank wall as on a face. But you often know in advance where the detail will be." Brian Chao's post on the same paper adds the reason this is timely. Image models increasingly start from a plan (boxes, masks, a layout written by a language model) before a pixel exists. If the plan already says where the hard parts are, why does the transformer still give the sky the same 4,096 tokens as the eyes?

Level-of-Token Diffusion, or LoT (arXiv 2610.05816), answers that literally. It comes from Stanford, first author Kiyohiro Nakayama, with Federico Tombari of Google and Leonidas Guibas among the authors and Lior Yariv and Gordon Wetzstein as equal advisers. It tiles the token grid with rectangles of different sizes, fine where the plan says detail matters, and fine-tunes a pretrained DiT to denoise the short sequence that tiling produces. Brian's thread puts it at "2–3.5×" faster while keeping quality and control.

I expected token merging in a new coat, and was wrong. The tokens are fixed before sampling starts, so nothing is measured or merged at run time, and the whole trick rests on one piece of algebra borrowed from the same group's asymmetric flow paper, which lets an 8×8 token look at 128 numbers out of 8,192 and still return a full-resolution velocity. The second thing that surprised me came from the speed tables. On images LoT never runs as fast as its token cut suggests. On video it runs faster. That difference tells you where to use it.

There is no code and no weights. The project page is a static site whose repository (georgenakayama/lotdiffusion on GitHub) holds the demos, real token layouts as JSON and a timing file, but no model. So everything below is checked against the paper, its appendix, the project page's data files and the two base models' published configs.

Top: a Level-of-Token layout on the left, a grid where detailed regions such as plant tendrils, a spiral and letters are tiled with small pink and peach tokens and the plain background with large pale-blue tokens, labelled 2.16x fewer tokens; on the right the generated image of plants arranged as a spiral and the words LEVEL-OF-TOKEN DIFFUSION, labelled 2.5x speedup. Bottom: applications. A depth map of a sleeping cat with its layout (2.00x fewer tokens, 1.58x speedup), a bounding-box living-room layout (2.39x fewer tokens, 1.78x speedup), and three video rows: a night driving scene from semantic masks (1.79x fewer tokens, 2.21x speedup), a toy train from 3D assets (1.49x, 1.75x) and a robot arm from a bounding box (1.50x, 1.76x).
LoT allocates tokens from a layout before generation. Note the pairs under each example: the two small images speed up less than their token cut, while the large teaser image and the three videos speed up more (LoT paper, Figure 1).

A layout is a tiling of the token grid

Start from what a pretrained DiT does. FLUX.2 klein encodes a 1024×1024 image to a 128×128 latent with 32 channels, then patchifies 2×2 latent pixels into one token. Its transformer/config.json says in_channels: 128, which is those 2×2×32 numbers, and the hidden width is 24 heads of 128, so 3,072. The DiT sees a 64×64 grid, 4,096 tokens, each a 128-vector. Every one of them goes through every block at every step.

LoT keeps that grid as the finest level and lets a token cover a rectangle of it. Formally a layout is a partition of the H×WH \times W token grid into LL disjoint axis-aligned rectangles

Ri=[ui, ui+eh,i)×[vi, vi+ew,i),R_i = [u_i,\, u_i + e_{h,i}) \times [v_i,\, v_i + e_{w,i}),

where (ui,vi)(u_i, v_i) is the top-left corner and ei=(eh,i,ew,i)\boldsymbol{e}_i = (e_{h,i}, e_{w,i}) the extent. The image model supports eh,ew∈{1,2,4,8}e_h, e_w \in \{1, 2, 4, 8\}, which is sixteen shapes including thin ones like 1×8. The video model, built on Wan2.1, uses {1,2,4}\{1, 2, 4\} per spatial axis and never merges across time. The paper reports savings as token compression, HW/LHW/L for an image and THW/LTHW/L for a video.

Left, a uniform token layout: a 4 by 4 grid of blue tokens each covering 2 by 2 latent cells, flattened into the sequence x_U. Right, a Level-of-Token layout over the same area: one tall orange token covering a 4-wide by 8-tall block labelled R_i with its extents e_h,i and e_w,i and top-left corner p_i, four blue tokens of the original size, and one large red token covering a 4 by 4 block. The resulting sequence x-bar_P has just six tokens.
A uniform layout against a level-of-token layout of the same region: six tokens instead of sixteen, each a rectangle with a corner and an extent (LoT paper, Figure 2).

What does a real layout look like? The project page ships the exact tiling behind one of its demos, the "flowering letters" image, as a list of [x, y, w, h] rectangles on a 96×160 grid (a 1536×2560 image). I rendered it next to the generation. Five rows of the word LOT get progressively more detailed from top to bottom, from flat paint to dense jasmine, and the layout follows: 554 tokens in the top fifth, 1,122 in the bottom fifth, against 3,072 for a dense fifth.

Left, a token layout on a tall 96 by 160 grid colour-coded by token area: large dark-blue 8 by 8 and mid-blue 4 by 4 tokens fill the background between letters; the five rows of L O T letters are outlined in small red 1 by 1 and orange and cream 2-cell and 4-cell tokens, and the bottom two rows are almost entirely red 1 by 1 tokens. Right, the generated image: five rows of green letters L O T on an ivory background, plain painted in the first row, polka-dotted in the second, leafy in the third, with vines in the fourth and covered in white jasmine flowers in the fifth.
A real LoT layout and its generation. 4,120 tokens against 15,360 for the dense grid; red is 1×1, orange two cells, cream four, light blue eight, mid blue 4×4 and dark blue 8×8. I drew the layout from the rectangle list in the project page's layout-inspector-data.js; the image is the project's own output (LoT project page).

Counting that list was useful. 4,120 tokens tile all 15,360 cells, a 3.73× compression. The 1×1 tokens are 69% of the sequence but cover only 18% of the image. The 4×4 and 8×8 tokens are 9% of the sequence and cover 61%. That asymmetry is what you are buying: the attention and the MLPs see mostly fine tokens, because most of the image is easy and costs almost nothing. It also contains 162 tokens of 2×1, 152 of 1×2, and a few 8×1 and 1×8 strips, so this is quadtree-like rather than a quadtree. The texture-variance rule coarsens one axis at a time along edges.

Where the layout comes from

The model never sees the signal that made the layout. It gets the text prompt and the rectangles, nothing else. Appendix A.3 describes four ways to make the rectangles, and they share one move: turn the signal into "how fine must this spot be", then tile.

Boxes and semantic masks assign a level ℓ∈{0,1,2,3}\ell \in \{0,1,2,3\} to each region, mapped to an extent 23−ℓ2^{3-\ell} (8, 4, 2, 1). The background gets a coarser level than any selected region, and where regions overlap the finer one wins. Texture variance adapts the variable-rate-shading rule from real-time graphics: a region is coarsened along a direction when the luminance error from doing so stays under a tolerance s(ER[I]+a)s(\mathbb{E}_R[I] + a), which is why it produces strips. Depth goes through a circle-of-confusion model, r(q)=K ∣d(q)−1−df−1∣r(\boldsymbol{q}) = K\,|d(\boldsymbol{q})^{-1} - d_f^{-1}|, so pixels far from the focal plane get large blur radii and large tokens. A region only merges when its smallest blur radius allows it, so one in-focus pixel keeps its block fine.

The fourth source is the one the agent demos use, and the simplest to reason about. Paint a detail map with four brush colours scoring 0, 0.3, 0.6 and 1, average it onto the token lattice, multiply by a gain, clip to [0,1][0,1]. Then start from 8×8 blocks and split a block of extent bb into four whenever its brightest cell clears a threshold:

max⁡q∈RD(q)≥tb,t8≤t4≤t2.\max_{\boldsymbol{q} \in R} D(\boldsymbol{q}) \geq t_b, \qquad t_8 \leq t_4 \leq t_2 .

That rule, with the real FLUX.2 lattice, is the widget below. The scenes are toys I drew; the split rule, the RoPE centre and the head sizes are the paper's and the model's. The paper doesn't publish its thresholds, so I picked 0.2, 0.45 and 0.75.

quadtree token layout · 64 x 64 lattice, 1024 pxsplit rule: paper Eq. 23 · thresholds mine

a subject box (0.6) with a finer box for the face (1.0); background 0.3

detail field D(q) after gain
1x1: 1196
2x2: 37
4x4: 172
8x8: 0
tokens in the DiT sequence
1405 of 4,096
compression 4096 / L
2.92x
speed, fit to paper's 4B timings
~2.03x
Click a token to see its position, shape features and head sizes.

Blocks start at 8x8 and split while the brightest cell under them clears the threshold for their size (0.2, 0.45, 0.75 here). The gain multiplies the whole detail field, so it moves every block at once, which is how the paper exposes the budget for painted maps. Real layouts from texture variance also contain rectangles such as 2x1 and 8x1; this quadtree only makes squares.

Two things stand out once you drag the gain. First, the budget is not a number you type. It is whatever the thresholds and the gain produce, which the paper says outright: the controls "let the user adjust the resulting token count without requiring an exact token budget". Second, the count moves in jumps. A whole region crosses a threshold at once and the sequence length snaps from 772 to 1,954 tokens on my painted scene. A real interface would want a search over the gain to land near a target, and the paper's evaluation instead sweeps VRS thresholds and reports whatever budgets come out (1.52× to 2.94× in the main tables, up to about 8× in the appendix sweep).

Eight pairs of LoT layout and generation in four columns: bounding box, semantic mask, texture variance and depth. Top row is a deer in a forest, bottom row a potted plant. Captions under each pair give compression and speedup: bounding box 1.73x and about 1.54x, 1.62x and 1.47x; semantic mask 2.18x and 1.81x, 2.88x and 2.18x; texture variance 1.36x and 1.28x, 2.00x and 1.71x; depth 3.11x and 2.29x, 2.00x and 1.71x.
One model, four layout sources. Every pair shows the compression and the speedup it bought (LoT paper, Figure 4).

Teaching a pretrained DiT to eat rectangles

Feeding a DiT a sequence of mixed-size tokens breaks three things at once. The input projection expects 128 numbers per token and an 8×8 token has 8,192. The output head emits one token's worth of velocity and the sampler needs a full-resolution velocity for every latent pixel, because the state being integrated and the VAE decoder both live on the full grid. And RoPE gives each token a grid position but says nothing about how big it is. LoT fixes each with as little new machinery as it can.

Pipeline. On the left the noisy full-resolution grid x_t, partitioned into an orange tall token, four blue small tokens and a red square token. A Patchify box with A transpose of e_i projects each region to one token, giving the short sequence x-bar_t,P; an Embed box takes t, c and the shape features s(e_i) and feeds g_phi into the patchify step and t, c into the DiT blocks. N DiT blocks process the sequence into hidden states h_i, which are routed to extent-specific heads h_(4,2), h_(1,1) and h_(2,2). The heads produce the asymmetric velocity u-hat_A on the full grid, which a formula on the right converts into the full-rank velocity: u-hat equals P times u-hat_A plus (I minus P) times (x_t plus u-hat_A) over sigma_t.
LoT's forward pass: project each region to one token, run the DiT on the short sequence with shape conditioning, decode each token with the head for its extent, then recover the full-rank velocity patch by patch (LoT paper, Figure 3).

The patch lift

For every extent e\boldsymbol{e} the authors fit a semi-orthonormal matrix Ae∈RDe×DA_{\boldsymbol{e}} \in \mathbb{R}^{D_{\boldsymbol{e}} \times D}, with De=ehewDD_{\boldsymbol{e}} = e_h e_w D. For an 8×8 token on FLUX.2 that is 8,192 by 128. A coarse token is the projection of the dense patch underneath it,

xˉt,Ri=Aei⊤xt(i)∈RD.\bar{\boldsymbol{x}}_{t,R_i} = A_{\boldsymbol{e}_i}^{\top} \boldsymbol{x}_t^{(i)} \in \mathbb{R}^{D}.

The fit is orthogonal Procrustes between dense patches and "extent-matched multi-scale VAE tokens" (Eq. 14). Read that slowly: the target is what the pretrained VAE plus patchify produce when the image is encoded at the coarser scale. So Ae⊤A_{\boldsymbol{e}}^\top maps a block of fine latent onto something shaped like a lower-resolution latent token, which the pretrained input layer has seen before. The extent's input head starts as the pretrained input projection composed with Ae⊤A_{\boldsymbol{e}}^\top, so on day one a coarse token embeds like a low-resolution one. That initialisation is where most of the pretrained prior survives.

A scalar ses_{\boldsymbol{e}} per extent fixes the magnitude mismatch (projection does not preserve token norms). AsymFlow handled this by rescaling the timestep. LoT can't, because a sample with mixed extents would then need a different timestep per token while the DiT was trained with one shared tt. So the authors divide each clean patch by its ses_{\boldsymbol{e}} before adding noise and multiply back before decoding (Eq. 16). Every token keeps the same tt and the projected noise stays standard Gaussian. Small decision, and the right one: it keeps the conditioning path exactly as pretrained.

Why a coarse token can still return a full velocity

This is the part I had to work through on paper. Flow matching here uses xt=(1−t) x0+t ε\boldsymbol{x}_t = (1-t)\,\boldsymbol{x}_0 + t\,\boldsymbol{\varepsilon} and the target velocity u=ε−x0\boldsymbol{u} = \boldsymbol{\varepsilon} - \boldsymbol{x}_0, where x0\boldsymbol{x}_0 is clean data. A coarse token only shows the network A⊤xtA^\top \boldsymbol{x}_t, 128 of 8,192 dimensions. It cannot possibly know the noise in the other 8,064.

Asymmetric flow (arXiv 2605.12964, Chen, Ackermann, Kim, Wetzstein and Guibas) changes the target so it doesn't have to. With the projector Pe=AeAe⊤P_{\boldsymbol{e}} = A_{\boldsymbol{e}} A_{\boldsymbol{e}}^\top, the network predicts

uA(i)=Peiε(i)−x0(i),\boldsymbol{u}_{\mathrm{A}}^{(i)} = P_{\boldsymbol{e}_i}\boldsymbol{\varepsilon}^{(i)} - \boldsymbol{x}_0^{(i)},

the clean patch at full resolution, but the noise only inside the subspace the token can see. The full velocity comes back without the network (Eq. 9):

u^(i)=Peiu^A(i)+(I−Pei) xt(i)+u^A(i)σt.\hat{\boldsymbol{u}}^{(i)} = P_{\boldsymbol{e}_i}\hat{\boldsymbol{u}}_{\mathrm{A}}^{(i)} + (I - P_{\boldsymbol{e}_i})\,\frac{\boldsymbol{x}_t^{(i)} + \hat{\boldsymbol{u}}_{\mathrm{A}}^{(i)}}{\sigma_t}.

Check it with the exact target. PuA=Pε−Px0=PuP\boldsymbol{u}_{\mathrm{A}} = P\boldsymbol{\varepsilon} - P\boldsymbol{x}_0 = P\boldsymbol{u}. For the other part, xt+uA=(1−t)x0+tε+Pε−x0\boldsymbol{x}_t + \boldsymbol{u}_{\mathrm{A}} = (1-t)\boldsymbol{x}_0 + t\boldsymbol{\varepsilon} + P\boldsymbol{\varepsilon} - \boldsymbol{x}_0, and applying (I−P)(I-P) kills the PεP\boldsymbol{\varepsilon} term and leaves t (I−P)(ε−x0)t\,(I-P)(\boldsymbol{\varepsilon} - \boldsymbol{x}_0). Divide by σt=t\sigma_t = t and you get (I−P)u(I-P)\boldsymbol{u}. The two pieces add up to u\boldsymbol{u}. The unseen noise is never predicted; it is read off the state the sampler already holds, once the clean patch is known.

So the job of a coarse token is to predict a full-resolution clean patch from one hidden vector. Its output head is linear, he:R3072→RDeh_{\boldsymbol{e}}: \mathbb{R}^{3072} \to \mathbb{R}^{D_{\boldsymbol{e}}}, initialised as AeWoutA_{\boldsymbol{e}} W_{\mathrm{out}}. For an 8×8 token that maps 3,072 numbers to 8,192, so the clean patches it can express live in a subspace of at most 3,072 dimensions. My reading is that this is where the smoothness of coarse regions comes from: the paper describes LoT "favoring smoother or defocused content in coarser-token regions", and a rank-limited linear decoder from one token can only paint so much texture into 64 latent positions. The paper doesn't make this argument; it is mine.

The ablation says how much the asymmetric target carries. Replace the recovery with heads trained to output the full-resolution flow directly ("direct HR prediction"), keep everything else, and FID goes from 13.80 to 26.49 at 1.52× compression and from 13.35 to 68.51 at 2.94×. Nothing else in the paper moves the numbers that much.

Position and size

Each token keeps the pretrained axial RoPE, indexed at its centre on the finest grid (Eq. 13):

ci=(ui+eh,i−12,  vi+ew,i−12).\boldsymbol{c}_i = \Big(u_i + \tfrac{e_{h,i}-1}{2},\; v_i + \tfrac{e_{w,i}-1}{2}\Big).

An 8×8 token at the corner sits at (3.5,3.5)(3.5, 3.5); a 1×1 token keeps its integer position, so a layout of all unit tokens is the pretrained model exactly. Fractional positions are free with RoPE, since the rotation is a continuous function of position. If you want the rotation itself from the ground up, the RoPE page covers it. The paper's point is that this is query-independent: unlike the multi-scale RoPE variants in Foveated Diffusion and similar work, no token needs a position map that depends on who is asking, so the attention stays a stock dense kernel.

Size goes in separately: four numbers (log⁡2eh,log⁡2ew,log⁡2ehew,log⁡2(eh/ew))\big(\log_2 e_h, \log_2 e_w, \log_2 e_h e_w, \log_2 (e_h/e_w)\big) through a zero-initialised MLP, added to the token embedding. Zero init means it starts as a no-op. It matters less than I expected. Turning it off costs FID 13.80 to 14.58 at 1.52× compression and 13.35 to 14.93 at 2.94×, small next to the other ablations. The extent-specific input and output heads already tell the network a lot about size.

None of the code is public, so I wrote the forward pass out from Eqs. 7 to 13 to make the shapes concrete. It is my sketch, not the authors' code:

# my sketch of one LoT denoising call (not released code); FLUX.2 klein 4B shapes
# x_t: (H0, W0, 32) noisy latent in the scaled space; layout: list of (u, v, eh, ew)
tokens, pos, shape = [], [], []
for (u, v, eh, ew) in layout:
    patch = patchify(x_t)[u:u+eh, v:v+ew].reshape(-1)        # eh*ew*128 numbers
    z = A[(eh, ew)].T @ patch                                # 128, Eq. 7
    tokens.append(W_in[(eh, ew)] @ z + g_phi(s(eh, ew)))     # 3072, Eq. 12
    pos.append((u + (eh - 1) / 2, v + (ew - 1) / 2))         # RoPE centre, Eq. 13
h = dit_blocks(tokens, rope=pos, t=t, text=c)                # L tokens instead of 4,096
u_hat = empty_like(x_t)
for (u, v, eh, ew), hi in zip(layout, h):
    ua = head[(eh, ew)](hi)                                  # eh*ew*128, Eq. 10
    P = A[(eh, ew)] @ A[(eh, ew)].T
    xt = patchify(x_t)[u:u+eh, v:v+ew].reshape(-1)
    full = P @ ua + (I - P) @ (xt + ua) / max(sigma_t, 1e-6) # Eq. 9
    write_patch(u_hat, u, v, eh, ew, full)
x_next = x_t + dt * u_hat                                    # ODE step on the full grid

In practice you would batch the patches by extent and never build PP as an 8,192-square matrix: Py=A(A⊤y)P\boldsymbol{y} = A(A^\top \boldsymbol{y}) costs two thin matmuls. The point of writing it out is the loop structure. The transformer runs on LL tokens; everything before and after it is per-patch linear algebra on the full grid, and that part does not shrink.

Training is a LoRA, not a new model

The image model is FLUX.2 klein base 4B with its text encoder and VAE frozen. The per-extent input and output heads, the shape MLP and the final adaptive norm are trained in full; the attention, MLP and timestep layers get rank-256 LoRA adapters. 50,000 steps at batch 32, AdamW at 1e-4, on Aesthetic-Train-V2 at 1024×1024 with a 95/5 split, and each training image gets layouts from SAM 3 masks and boxes, texture variance and a depth estimator. The loss is a clean-data MSE weighted by 1/max⁡(σt,0.05)21/\max(\sigma_t, 0.05)^2, with power-function EMA (γ=7\gamma = 7) for evaluation.

A 9B variant trained separately on three million LAION images, 32 H100s, batch 256, and the qualitative demos use its 24,500-step EMA checkpoint. The video model is Wan2.1 T2V 14B, 4,000 steps at batch 32 on 81-frame 1280×720 clips from Vchitect's dataset, with the layout mix 50% masks or boxes, 25% texture variance, 25% depth of field. Its grids are capped at 3,072 tokens per latent time slice to avoid running out of memory. Wan's dense slice at 720p is 45×80 = 3,600 tokens, so every training clip was compressed at least a little; I'd like to know how the model behaves at full density, and the paper doesn't say.

The cheapness is the attraction. These are fine-tunes of open checkpoints, and the method touches the model in four places: the heads, a small MLP, the RoPE index and LoRA. If code appears, it would be a modest job to port to another flow DiT.

Where the time goes

Now the speed. The image numbers come from Table 1 plus the extra budget in Table 5, all on an H100 at 30 steps, timed from prompt encoding to VAE decode:

token compressionLoT timeLoT speedupratio
1.00× (dense control)5.77 s1.00×
1.52×4.37 s1.32×0.87
1.89×3.70 s1.56×0.82
2.41×3.18 s1.82×0.76
2.94×2.83 s2.04×0.70

The last column is speedup divided by compression, and it falls as compression rises. I fitted t(c)=a+b/ct(c) = a + b/c to the dense reference (5.7783 s) and the four LoT timings: a≈1.33a \approx 1.33 s, b≈4.49b \approx 4.49 s, with every residual under 0.09 s. Read aa as the work that doesn't shrink with the token count: text encoding, the VAE decode, the per-patch recovery on the full grid, launch overhead. It is about 23% of a dense image, and it caps the speedup near 5.8/1.33, about 4.4×, however coarse the layout gets.

The video table goes the other way. At 2×, 2.5× and 3× compression LoT-Wan2.1 runs 2.26×, 2.91× and 3.53× faster. Speedup beats compression at every budget, and so does every baseline's.

Sequence length explains both. Per token and per block, the linear layers cost a fixed amount while attention costs something proportional to the sequence length. My rough count from the configs, ignoring text tokens and fixed costs: on FLUX.2 klein 4B (width 3,072, MLP ratio 3) a token's linear layers cost about 189 MFLOPs per block and its attention about 50 MFLOPs at 4,096 tokens, so attention is roughly a fifth of the block and the block cost falls about linearly with the token count. On Wan2.1 14B (width 5,120, FFN 13,824), a 720p, 81-frame clip is 21×45×80 = 75,600 tokens; linear layers are about 493 MFLOPs per token and attention about 1,548. Attention is three quarters of the work, and it shrinks with the square of the token cut.

The project page's own timing file fits the same pattern. Its five bounding-box demos run the 9B model at 1536×1024, 6,144 dense tokens, 50 steps: 28.7 s dense against 11.5 to 14.9 s for LoT, which works out to 1.93× to 2.49× for 2.32× to 3.15× fewer tokens. The teaser image goes 2.5× faster for 2.16× fewer tokens, and Wetzstein's thread puts its dense count at 14,336. Longer sequences, better ratio.

speed and quality at matched token budgets · paper Tables 1, 2 and 5H100, layouts from texture variance
0.801.21.62.02.41.01.62.12.73.2token compression (x)speedup vs dense (x)LoT (ours): 1.523x compression, 1.323LoT (ours): 1.893x compression, 1.56LoT (ours): 2.406x compression, 1.818LoT (ours): 2.935x compression, 2.04ToMe-SD: 1.893x compression, 1.538ToMe-SD: 2.406x compression, 1.793ToMe-SD: 2.935x compression, 2.003DDiT: 1.893x compression, 1.34DDiT: 2.406x compression, 1.521DDiT: 2.935x compression, 1.321Foveated Diffusion: 1.893x compression, 1.04Foveated Diffusion: 2.406x compression, 1.1Foveated Diffusion: 2.935x compression, 1.180.162.95.78.4111.01.62.12.73.2token compression (x)HPSv3 (higher better)denseLoT (ours): 1.523x compression, 10.3477LoT (ours): 1.893x compression, 10.1211LoT (ours): 2.406x compression, 9.7533LoT (ours): 2.935x compression, 9.5522ToMe-SD: 1.893x compression, 5.3569ToMe-SD: 2.406x compression, 2.915ToMe-SD: 2.935x compression, 0.9253DDiT: 1.893x compression, 9.0052DDiT: 2.406x compression, 8.0549DDiT: 2.935x compression, 7.4075Foveated Diffusion: 1.893x compression, 9.0092Foveated Diffusion: 2.406x compression, 8.2674Foveated Diffusion: 2.935x compression, 7.8433
LoT (ours)ToMe-SDDDiTFoveated Diffusion

Left: the dashed diagonal is speedup equal to compression. Image points sit below it, video points above it. Right: the dashed horizontal line is the dense model fine-tuned on the same data. Every method at a given budget gets the same layout, derived from a reference image or a reference video.

The abstract promises "significant speedups determined by the layout's token budget". True, with a footnote that matters for deployment. At 1024² on a 4B model, half the tokens buys you about 1.6×. On long video, or images well past 1024², the same layout buys more than its token count. If I were picking where to try this first, it would be video.

Reading the comparison tables carefully

Table 1 is where LoT looks best, and it is worth knowing what it measures. All methods are evaluated on the 5,261 held-out prompts, and for each prompt the layout is derived by texture variance from that prompt's held-out reference image. The paper is upfront that this "is a layout-conditioned comparison: the layout provides reference-derived spatial detail allocation beyond the text prompt". The generator never sees the reference pixels, but it is told where the reference image has texture.

That explains the most striking row. The dense control, FLUX.2 4B fine-tuned on the same split for the same 20,000 steps, scores FID 16.75. LoT at 1.89× fewer tokens scores 13.36. A model with fewer tokens beats the dense model on fidelity to the reference set, and on pFID and TOPIQ too. Appendix Figure 14 sharpens it. LoT's FID is U-shaped in compression, higher near 1.2× than around 3×, where it bottoms out, then climbing again toward 8×. I read both as the layout carrying information about the test images rather than as coarse tokens making better pictures. On the learned human-preference scores, which don't look at the reference set, LoT falls steadily below the dense model as the budget shrinks: HPSv3 10.43 dense, 10.12, 9.75, 9.55; ImageReward 0.95, 0.88, 0.85, 0.83. Those scores are the honest cost curve, and it is a gentle one.

Against the other token-reduction methods the picture is cleaner, because they get the same layouts and the same 20,000-step fine-tune (except ToMe-SD, which is training-free). LoT has the best HPSv2.1, HPSv3, FID, pFID, TOPIQ and MUSIQ at every budget, and is the fastest at every budget, though only just ahead of ToMe-SD (1.56× against 1.54× at the first budget). ToMe-SD collapses on quality: HPSv3 0.93 at 2.94×. The only place LoT loses is ImageReward at the first budget, where DDiT's 0.885 edges its 0.883.

Two caveats on the baselines. DDiT and Foveated Diffusion were extended by the authors to support 8×8 tokens, so their numbers are for ports, not their original code. And Foveated Diffusion, the same group's earlier mixed-resolution method, runs only 1.04× to 1.18× faster than dense in this table despite the same token cuts. That looks to me like an implementation cost (its query-dependent position maps) more than a ceiling of the idea, but the paper doesn't analyse it.

Top, four line charts of image quality against compression ratio from 1 to about 8 (4,096 over actual tokens), comparing Ours (blue) with FLUX.2 token matched (orange). FID: ours starts near 15, dips to about 13 around 3x, rises to about 16 at 8x; the token-matched baseline sits between 18.5 and 20. pFID: ours rises from about 17 to 30, baseline from 22 to 32. MUSIQ: ours falls gently from 71 to 65, baseline steeply from 70 to 53. TOPIQ: ours 0.62 to 0.53, baseline 0.57 to 0.36. Bottom, ten qualitative pairs: small token-matched FLUX.2 images such as 592 by 592 with 1,369 tokens, shown at their real size on hatched padding, next to LoT 1024 by 1024 outputs with similar token counts such as 1,404 tokens.
Against simply generating at a lower resolution with the same token count and upsampling, LoT wins clearly on every curve. Note the U-shaped FID for LoT (LoT paper, Figure 14).

The most useful baseline in the whole paper is in that appendix figure, and it is the boring one: pretrained FLUX.2 run at a lower resolution with a token count matched per image, then bicubically upsampled. Against that, LoT's MUSIQ falls from about 71 to 65 across the sweep while the low-resolution baseline falls from 70 to 53. If your alternative to LoT is "render smaller and upscale", LoT wins and the gap widens with compression. I'd have liked MrFlow-style resolution climbing in the same chart, since it attacks the same budget from the time axis and the paper says the two are compatible.

(a) Image generation at a budget of 1,269 tokens: a box of coloured clay sticks. Full-resolution FLUX2 4B at 1.00x speedup is sharp. ToMe-SD at 2.10x is a blocky, smeared mess. DDiT at 1.90x is blurry and noisy. Foveated Diffusion at 1.63x shows distortions at token boundaries. Ours at 2.29x is sharp and detailed. (b) Video generation at a budget of 50,134 tokens: three frames of two women by a hut. Full-res Wan2.1 14B at 1.00x; ToMe-SD at 1.60x is noisy and corrupted; Foveated Diffusion at 1.59x and Ours at 1.62x are both coherent.
Same token budget, four methods. ToMe-SD and DDiT leave noisy low-resolution regions; Foveated Diffusion distorts near mixed-resolution boundaries (LoT paper, Figure 7).

The video table follows the same shape on the two VBench scores that track appearance. Imaging quality is 0.647 dense and 0.612, 0.585, 0.565 for LoT at 2×, 2.5× and 3×; Foveated Diffusion is lower at each, ToMe-SD much lower. On the temporal scores LoT is level with Foveated Diffusion and slightly below dense. The evaluation layouts here come from reference videos generated by a different model (MiniMax-H3), because the held-out clips "were mostly static". Fair enough, and a reminder that the layouts in every table are oracle-ish: they describe where detail ended up in some real or generated video, not a guess made from the prompt.

The layout also moves the subject

Appendix A.9 runs an experiment that changed how I think about the method. Take 2,096 COCO prompts with one clear subject. Build a layout with 1×1 tokens inside the subject's real box and 2×2 outside, then a second layout with the same box translated somewhere else. Same prompt, same noise, same token count. The subject follows the fine tokens: the share of the detected subject inside the target box is 61.17% for LoT-FLUX.2 4B against 45.38% for dense FLUX.2, and 66.83% against 44.94% at 9B.

The authors present this as controllability, and it is. It is also a warning. A LoT layout is not a neutral compute hint; the model has learned that fine tokens are where things go. If you feed it a layout that disagrees with the prompt, it will bend the picture toward the layout. In the agent and storyboard settings the paper is aimed at, where a planner draws the layout and the layout is the intent, that is a feature. As a drop-in accelerator for arbitrary prompts, you need a layout source you trust, and the paper does not test one that is predicted from the prompt alone.

Four video rows, each three frames of layout and generation. Bounding box: a robot arm over a drawer, 1x1 tokens on the arm and object, 2x2 on the table, 1.50x compression and 1.60x speedup. Semantic mask: a night driving scene, fine tokens on lanes and vehicles, 2x2 on road and buildings, 2.00x and 2.32x. Texture variance: a desert game scene with a running character, 1.94x and 2.24x. Depth: a cyclist on a road, fine tokens near the camera and subject, 1.61x and 1.78x.
Per-frame layouts for video from four sources. Every row speeds up more than its token cut (LoT paper, Figure 5).

Where it breaks

The paper's own limitation figure is the one I would have asked for. Sheet music is fine structure everywhere. At 3,189 tokens (1.28× compression) the staff lines hold; at 1,466 tokens (2.79×) they fragment and bend. There is nothing to save when every region is hard, and forcing a budget anyway breaks structure first. Thin, long, regular things like text lines, cables and fences are what I'd expect a rank-limited coarse token to struggle with too, though the paper only shows the sheet music.

Two tokenization and generation pairs for the prompt 'Sheet music with handwritten notations and dynamic markings'. Left, higher budget: 3,189 tokens, 1.28x compression, a mostly fine grid, and a clean page of sheet music. Right, lower budget: 1,466 tokens, 2.79x compression, many coarser grey blocks, and a page where the staff lines are broken and wavy.
Fine structure everywhere leaves little to save, and a forced lower budget breaks the staff lines first (LoT paper, Figure 22).

A few smaller things I could not reconcile. The text says LoT-FLUX.2 achieves "1.32–2.04×" speedup in Table 1, but Table 1's lowest LoT speedup is 1.560×; 1.32 is the 1.52× budget that appears only in the ablation table. The video range is given as "1.58–3.53×", and I could not find 1.58 in any video table; the lowest in Table 2 is 2.262×. Wetzstein's thread mentions "up to 4.6× when detail is more concentrated", which no table in the paper reports. The typography gallery on the project page labels its outputs "LoT-Flux.2-14B" and "Full-resolution Flux.2-14B", a size the paper never mentions; the 4B and 9B are the only models it describes. None of these change the story, but they are the kind of thing a code release would settle.

Would I use it

The core idea is right and overdue. Generation pipelines are growing a planning stage, and that stage already knows where detail goes; letting it set the compute is cleaner than having the DiT rediscover it every step through attention scores or similarity merges. LoT does it with almost no new architecture, and the asymmetric-flow recovery is the piece I would reuse first, because it answers the question every coarse-token method has to answer (how do you get a full-resolution velocity out of fewer tokens) with algebra instead of a decoder.

For 1024² images on a 4B model, the speed is real but modest, about 1.6× at half the tokens, and the fixed cost puts a ceiling near 4.4×. For long video it is better than linear. The quality cost on preference scores is small and smaller than any baseline's. The two things I would check before trusting it in a product are what happens with a layout predicted from the prompt alone, and how far the layout pulls content you didn't ask for. Neither is measurable until the weights are out. The neighbours on this site attack the same budget elsewhere: Sana shrinks the token grid with a deeper autoencoder, PixelUMM and PixelDense drop the VAE altogether, and Mixture-of-Depths is the language-model cousin where a router picks which tokens a block computes. For the DiT itself, start at the Diffusion Transformer page; FLUX 3 is where the FLUX line went next.

How I checked

I read the paper and its appendix from the arXiv HTML (v1, 5 October 2026), and the X threads by Brian Chao and Gordon Wetzstein through the fxtwitter mirror. I shallow-cloned the project page's repository, georgenakayama/lotdiffusion at commit 5edcc89, which holds no model code. From it I counted the flowering-letters layout in layout-inspector-data.js (4,120 rectangles covering 15,360 cells, per-extent counts and per-fifth counts), rendered that layout for the figure above, and read supplementary/bbox_timings.json and bbox_examples.json for the 9B timings and token counts. Model shapes come from the published configs: FLUX.2 klein base 4B (in_channels 128, 24 heads of 128, 5 double-stream and 20 single-stream blocks, MLP ratio 3) and Wan2.1 T2V 14B (width 5,120, FFN 13,824, 40 layers, patch 1×2×2, 16 latent channels).

The speed-to-compression ratios, the a+b/ca + b/c fit and its 4.4× ceiling, and the per-token FLOP split are my arithmetic on the paper's tables and those configs. The FLOP split ignores text tokens, the double-stream blocks' separate text weights and kernel efficiency, so treat it as a direction, not a prediction. I checked the Eq. 9 recovery by hand. The idea that rank-limited coarse heads explain smooth coarse regions is my reading, not the paper's. I could not run anything: no code or weights have been released.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Level-of-Token Diffusion: a DiT that is told where the detail goes", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026lotdiffusion,
  author = {Satyajit Ghana},
  title  = {Level-of-Token Diffusion: a DiT that is told where the detail goes},
  url    = {https://ai.thesatyajit.com/articles/lot-diffusion},
  year   = {2026}
}
share