~/satyajit

T3-Video: 685,440 tokens per 4K clip, and the window attention that makes them affordable

mdjsonmcp

2026-10-08 · 23 min · video-generation · diffusion · diffusion-transformers · attention · sparse-attention · fine-tuning

Why read this

Essentialtop 10%

Recounts T3-Video's MAC table from code, maps every layer's windows, and finds a duplicated attention pass and a misread theory-vs-practice gap.

  • Original, source-checked analysis
  • Interactive explanations
  • A lasting reference

Image & video generationNeeds datacenter GPUsApache-2.0Practitioner paper

How this was scored
Is it new?
2 of 3: A real new idea, method or capability
Can I trust it?
3 of 3: Reproduces the headline result, or shows from primary files it is wrong
Can I run it?
2 of 3: Open code or weights with real limits
Will I understand it?
3 of 3: Mechanism carried by interactives built from real code or data
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
2 of 3: A reference for a year or more
Does it affect many?
1 of 3: A specialist community
Only here?
3 of 3: The only place this analysis exists

Score 81 of 100, ranked 14 of 476 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

A post from @HuggingApps came with a claim I wanted to take apart: "native 4K video, from a 1.3B model. T3-Video retrofits Wan2.1-1.3B with window attention so it generates directly in 3840×2176, no upscaler." The demo zooms into a canyon until single bushes fill the screen.

The Space links three things: the paper, Transform Trained Transformer for Accelerating Native 4K Video Generation (Jiangning Zhang and eleven co-authors; the README's citation lists ICML 2026), the code at github.com/zhangzjn/T3-Video, and the weights at APRIL-AIGC/T3-Video. I read the paper end to end, cloned the code, read the safetensors headers and pulled the Space's own app.py. My question was simple: a 1.3B model trained at 480p is now being asked to denoise 685,440 tokens per step. How is that not a week of GPU time?

The answer turns out to be clean and a little surprising. The weights keep their exact shapes. What changes is which tokens each attention call is allowed to see. And once I recomputed the paper's cost table from the code, two things fell out that the paper does not say: one layer in every five runs the same attention twice, and the "gap between theoretical and actual speedup" it apologises for mostly is not there.

APRIL-AIGC/T3-Video@b38f51b · snapshot 2026-10-08
repo size
12.84 GB
task
text-to-video
library
diffusers
license
apache-2.0
safetensors
2 shards
largest file
10.00 GB
files
7
downloads
101
likes
20
videot2vi2v

repo last modified 2026-10-05

zhangzjn/T3-Video@09454d4 · snapshot 2026-10-08
tracked files
242
license
Apache-2.0
branch
main
tests
none found
source
2.9 MB
commit date
2026-10-05
source by language
Python2.9 MB(184)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-08 at 09454d4 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow

shallow clone: counts describe the pinned tree, not the history

Counting the tokens

Start with what the transformer actually receives. Wan2.1's VAE compresses 8× in each spatial direction (upsampling_factor = 8, wan_video_vae.py:1077) and 4× in time while keeping the first frame on its own, so 81 frames become (81 - 1) // 4 + 1 = 21 latent frames (wan_video_new.py:495). The DiT then patches every 2×2 latent pixels into one token (patch size [1, 2, 2], wan_video_dit.py:640). One token is a 16×16 pixel square over four frames.

At 3840×2176 that is 240 tokens across, 136 down, 21 deep:

L=21×136×240=685,440 tokensL = 21 \times 136 \times 240 = 685{,}440 \text{ tokens}

The Space's header says "685k tokens", which is this number. For comparison, the same 81 frames at 832×480 give 32,760 tokens. Going to 4K multiplies the token count by about 21. Full self-attention costs grow with the square of that, so the attention bill goes up roughly 440-fold.

The paper writes the cost of one transformer layer in multiply-accumulates, ignoring biases, with CC the width and CffnC_{ffn} the FFN width:

MACsfull=4LC2+2L2C+2LC Cffn\text{MACs}_{full} = 4LC^2 + 2L^2C + 2LC\,C_{ffn}

The first term is the Q, K, V and output projections, the last is the FFN, and the middle one is attention itself: QK⊤QK^\top and the weighted sum over VV, each L2CL^2C. Wan2.1-1.3B has C=1536C = 1536, Cffn=8960C_{ffn} = 8960 and 30 layers (wan_video_dit.py:642-649). Plug in L=685,440L = 685{,}440 and the attention term alone is 43,299 trillion MACs per forward pass, against 858 trillion for everything else combined. At 4K, attention is 98% of the model.

Memory is the other wall, and it is subtler than it looks. If you wrote the attention scores out, one layer's 12 heads at bf16 would be 685,4402×12×2685{,}440^2 \times 12 \times 2 bytes, about 11.3 TB. Nobody does that: FlashAttention streams the scores through on-chip memory and never stores them. What you actually hold is activations. One L×1536L \times 1536 tensor in bf16 is 2.1 GB, and the FFN's L×8960L \times 8960 hidden state is 12.3 GB. Activations are why the paper reports 59.5 GB to run inference at 4K even with windowed attention, and why the Space had to patch the code to fit a 48 GB budget (more on that below).

tokens and attention cost · Wan2.1-1.3B vs T3-Videoone DiT forward, MACs
resolution
latent grid T x H x W
21 x 136 x 240
tokens L
685,440
attention saving
43.0x
whole-DiT saving
23.7x
full attention · 30 layers x 2·L²·C43,299T
T3 window attention · both branches, as shipped1,007T
projections + FFN + rest · identical in both models857.8T
bars on a log scale from 1T
paper, Tables 1 and 2, at 4K x 81 frames
attention 43,299.3T vs 1,006.8T (43.0x) · whole DiT 44,157.1T vs 1,864.5T (23.7x)
measured DiT time, 50 steps on one H20: 39,661.7s vs 1,857.4s (21.4x)
attention scores, one layer, 12 heads, bf16, if written out
full: 11.3 TB · T3 small-window layer: 44.2 GB
what flash attention leaves you holding, bf16
one L x 1536 tensor: 2.1 GB · FFN hidden: 12.3 GB

Wan2.1's VAE divides each side by 8 and time by 4, and the DiT patches 2 x 2 more, so one token stands for a 16 x 16 pixel square over four frames. At this setting that is 685,440 tokens. Full attention costs 43,299T MACs per forward pass; T3's windows cost 1,007T. The projections and FFN, 857.8T, do not shrink at all, which is why the whole-DiT saving is smaller than the attention saving.

Drag the frame count and switch resolutions and the shape of the problem is plain. The projection and FFN bars move linearly. Full attention moves quadratically and swamps them by 4K. The T3 bar moves linearly too, because its windows stay a fixed size while the video grows.

What happens if you just ask Wan for 4K

Nothing about Wan2.1 stops you from setting height=2176, width=3840. RoPE extends, the convolutions do not care. The paper did it.

Four 4K text-to-video results side by side. Wan2.1-T2V-1.3B at 4K produces a flat teal frame and is labelled 33 hours. HunyuanVideo produces a blurred, smeared scene and is labelled 45 hours. UltraGen produces a plausible elephant and a beach scene and is labelled 7 hours. T3-Video produces a sharp elephant and three people holding cake on a beach and is labelled 1 hour. An inset bar chart compares attention MACs and DiT latency for Wan2.1-1.3B at 480P, 720P, 1080P and 4K, rising from 99T and 131 s to 43299T and 39661 s on a log scale.
Running the official models directly at 4K, against UltraGen and T3-Video. The inset plots Wan2.1-1.3B's attention MACs (blue) and measured DiT latency (orange) by resolution, log scale. Efficiency measured on one H20 with FlashAttention-2 (T3-Video paper, Figure 1).

Wan2.1-1.3B at 4K renders a flat teal field. HunyuanVideo smears. Both take a day or more. The model was never trained to place things at these positions, and the attention distribution over 685,440 keys is nothing like the one it saw over 32,760. Brute force is both slow and wrong.

One number in this figure does not match the paper's own table, and I will come back to it: the "33 Hours" label for Wan.

Two windows per query

The obvious fix is window attention: cut the token grid into blocks and let each token attend only within its block. If a block holds LbL_b tokens, the attention term becomes 2LLbC2LL_bC, linear in LL. The paper's Figure 2 shows why nobody can stop there. With purely local windows, each block draws its own little scene, and you get a mosaic of disconnected cats. Purely strided windows, where each block takes every n-th token across the whole frame, produce noise.

Four 720P generations in a two by two grid. (a) Close windows only, with and without fine-tuning: a mosaic of tiles, each a separate small scene of a cat by a log pile. (b) Remote windows only: coloured noise. (c) T3 without fine-tuning: blocky coloured noise. (d) T3 with fine-tuning: one coherent image of a cat sitting by a tree trunk and a brick wall.
Close windows alone tile the frame into unrelated scenes; remote windows alone give noise; T3 combines both and needs fine-tuning before it produces a coherent image. 720P, 4×4 blocks (T3-Video paper, Figure 2).

T3's answer is to run both at once and average. In every self-attention layer, each token takes part in two attention calls of the same size. The close call covers a contiguous block around it. The remote call covers a dilated grid: same number of tokens, but sampled every few rows and columns across the entire frame. Both use the layer's existing WQW_Q, WKW_K, WVW_V and WOW_O. Nothing new is added to the model.

Diagram of a grid of latent tokens with T, H and W axes. Token F(1,1) sends orange arrows to its immediate neighbours inside a shaded orange block, blue arrows to tokens a medium stride away, and green arrows to tokens at the far edges of the grid. A legend notes that all three sets of arrows share parameters.
One token exchanging information at several scales at once, with the same attention parameters at every scale. The released model uses two scales: the adjacent block and the widest stride (T3-Video paper, Figure 3).

The code is a pair of einops rearranges that differ only in the order of two axes. This is the whole mechanism, from diffsynth/models/wan_video_dit.py:

diffsynth/models/wan_video_dit.py:242-272
q_remote = rearrange(q, 'b (w_t n_t) (w_h n_h) (w_w n_w) d -> (b n_t n_h n_w) (w_t w_h w_w) d', ...)
x_remote = self.attn(q_remote, k_remote, v_remote)
...
q_close = rearrange(q, 'b (n_t w_t) (n_h w_h) (n_w w_w) d -> (b n_t n_h n_w) (w_t w_h w_w) d', ...)
x_close = self.attn(q_close, k_close, v_close)
...
x = (x_remote + x_close) / 2

Read (n_h w_h) as "split the height into n_h blocks of w_h rows" and you get contiguous blocks. Read (w_h n_h) as "take every n_h-th row" and you get a strided set. Each becomes a batch of short sequences fed to the ordinary flash-attention call. Because the windows are just a reshape into the batch dimension, the attention kernel itself is untouched, which is what the paper means by a "one-line" replacement and why it can keep using stock FlashAttention-2.

I like this design. It is cheap to reason about: within one layer every token sees a sparse sample of the whole video through the remote grid, and full-resolution context through the close block. It is the dilated-plus-local trick from image backbones, which the paper credits to its own EMOv2 work, applied without inventing a module.

Five layer types, cycled six times

The paper says the blocking scheme varies by layer, grouped in fives, and leaves the numbers to the code. They are hard-coded in SelfAttention.__init__ (wan_video_dit.py:199-212), keyed on the layer index modulo 5. (n_t, n_h, n_w) is the number of blocks along each axis of the 21×136×240 grid:

layer index mod 5blocks (n_t, n_h, n_w)close window (frames × rows × cols)in pixelstokens per window
021, 1, 11 × 136 × 240one whole frame32,640
11, 17, 3021 × 8 × 8128 × 128, all frames1,344
21, 8, 4021 × 17 × 6272 × 96, all frames2,142
33, 17, 87 × 8 × 30128 × 480, 7 latent frames1,680
47, 8, 63 × 17 × 40272 × 640, 3 latent frames2,040

Two of the paper's ideas are visible here. The block boundaries move from layer to layer: an 8-row boundary in type 1, 17-row in type 2, 30-column in type 3, so a token cut off from its neighbour by one layer's grid is in the same block in the next. Swin gets the same effect by shifting windows; T3 gets it by changing their shape. And types 1 and 2 span all 21 latent frames while type 0 spans a whole frame, which is the paper's "axis-preserving full attention": every few layers some axis is attended in full. The paper's ablation found that this mix of shapes matters: configurations pushed toward only large or only small ratios scored 67.14 and 68.69 VQA at 720p against 69.37 for the shipped one.

what one token attends to · 4K latent frame, 136 x 240 tokensfrom wan_video_dit.py
latent frames in each window (query frame 10)
close 21 x 8 x 8 · remote stride 1 x 17 x 30 · 1,344 tokens per window · layers 1, 6, 11, 16, 21, 26

The close branch sees a solid block around the query. The remote branch sees the same number of tokens, one every 17 rows and 30 columns, so it spans the whole frame at low density. Both use the same Q, K and V projections, and their outputs are averaged. Click or drag on the frame to move the query.

Pick a layer type above and click around the frame. In type 1 the close block is a tiny 8×8 square, but the remote set is 64 single tokens spread across the whole 4K frame, 17 rows and 30 columns apart, through all 21 latent frames. The query can see the far corner of the frame, just sparsely.

The layer that does its work twice

Click type 0 in that widget and the remote and close windows coincide. I did not draw it that way for convenience.

With n_t = 21, w_t = 1, and n_h = n_w = 1, the two rearrange patterns (w_t n_t) and (n_t w_t) index the same frames, and (w_h n_h) with n_h = 1 is the same as (n_h w_h). I reproduced both permutations in NumPy on an index grid of the real shape: for type 0 the two branches gather identical sets of tokens, for types 1 to 4 they do not. So in layers 0, 5, 10, 15, 20 and 25 the model runs full-frame attention over 32,640 tokens, then runs it again on the same inputs, then averages two identical outputs.

This is not a small inefficiency. Those per-frame windows are by far the largest in the network. Summing 2LLbC2LL_bC over both branches and all 30 layers, I get 1,006.8 trillion MACs of attention per forward pass, exactly the figure in the paper's Table 1, and 82% of it comes from the type 0 layers. Dropping the duplicate branch cuts attention to 594.5 trillion MACs, which takes the reduction against full attention from 43.0× to 72.8×, and the whole DiT forward from 1,864.5 to 1,452.2 trillion MACs, about 22% less. The output would be bit-for-bit the same: the mean of two identical tensors is the tensor.

I have not run this. If the DiT stays compute-bound, as the measurements below suggest it is, a one-line if around the second branch should take something like a fifth off every 4K generation. The released weights would not need to change.

What was actually trained

A model whose attention pattern changes this much cannot just be loaded and run. Figure 2(c) shows T3 on unchanged Wan weights: noise. So I wanted to know exactly what the released checkpoint is.

The safetensors header of T3-Video-Wan2.1-T2V-1.3B.safetensors lists 825 tensors and 1,418,996,800 parameters, the same names and shapes as Wan2.1-T2V-1.3B's diffusion_pytorch_model.safetensors, with no LoRA tensors and nothing added. Wan ships in F32; T3 ships in BF16. Paper Table 1 says the same: 1419.0M parameters for both. This is a full fine-tune of every weight.

I then range-read four matrices from each of the 30 blocks in both files and compared them. Every tensor moved. The relative change ∥WT3−WWan∥/∥WWan∥\|W_{T3} - W_{Wan}\| / \|W_{Wan}\| averages 6.5% for the self-attention query projection and 5.9% for the key projection, but only 2.9% for the cross-attention query and 2.7% for the FFN output. The training concentrated on the part whose job changed: self-attention moved about twice as far as anything else. The biggest moves are at the ends of the stack, 10.2% for block 0's query and 10.9% for block 28's. I expected the type 0 layers, which keep full spatial attention, to move least. They do not. The layer type makes no visible difference to how far the weights travelled.

The recipe in the paper has three stages:

  1. Convert Wan2.1 to T3 attention and fully fine-tune at 720×1280 on UltraVideo, 42,184 high-resolution clips with long captions, from which 120 were held out as 4K-VBench. AdamW, learning rate 2e-5, batch size 64, DeepSpeed ZeRO-2, on H20s; the paper says 64 GPUs, against 128 H20s for UltraWan and UltraGen.
  2. Fully fine-tune that model again at 2176×3840 with 81 frames. The paper says this progressive route converges in 500 iterations; its training-progress figure says results are "satisfactory" by 2,000, and the general training details say 5K iterations. I could not pin down which run produced the released file.
  3. Optionally, a LoRA (rank 64, learning rate 1e-4) instead of the 4K full fine-tune. It scores slightly lower and is not what is released.

Two figures explain why the stages exist. Fine-tuning at 720p does not transfer to other resolutions for free:

Generations from the 720P-trained T3-Video model run at higher resolutions. Text-to-video outputs break into blotchy, incoherent textures as resolution rises; image-to-video outputs keep the first frame's layout but degrade.
A T3-Video model fine-tuned at 720P degrades when run at higher resolutions; the image prior in I2V softens the failure. Hence a second fine-tune at 4K (T3-Video paper, Figure 5).

And going straight from Wan's weights to T3 attention with only a LoRA fails, at ranks 32, 64 and 128 alike:

Rows of frames from direct LoRA fine-tuning of T3-Video from the official Wan weights. The text-to-video model produces noise; the image-to-video model keeps the first frame but shows blocky artefacts in later frames.
LoRA directly from the official weights cannot bridge the change in attention pattern; T2V fails outright and I2V turns blocky. A full fine-tune at low resolution comes first (T3-Video paper, Figure 7).

That has a practical consequence. The released window shapes are constants in __init__, the assert self.H % self.n_h == 0 checks enforce them, and the Space README says the checkpoint "only supports its native training shape: 2176×3840, 81 frames". You cannot ask this model for a 1080p clip or a vertical video. A different shape means different window constants and, per Figure 5, more fine-tuning. Note also that 2176 is not 2160: the height was picked because 136 divides by 17 and by 8, so you get 16 extra rows to crop.

An hour, not a day

Table 2 is where the paper's speed claims live. It took me a minute to see how its columns fit together. "DiT (50)" is the time for 50 transformer forwards, and the latency column is twice that plus the VAE decoder, because classifier-free guidance runs a conditional and an unconditional pass each step. For T3 at 4K: 2×1,857.4+451.0=4,165.82 \times 1{,}857.4 + 451.0 = 4{,}165.8 seconds. For stock Wan2.1 at 4K: 2×39,661.7+451.0=79,774.42 \times 39{,}661.7 + 451.0 = 79{,}774.4 seconds. The arithmetic holds in every row.

So on one H20, 50 steps:

The deployment numbers come with a catch. The weights repo holds two files, the T2V-1.3B and T2V-5B DiTs. There is no distilled checkpoint and no eVAE, and the training code is still an unchecked box in the README. The 166.8-second 4K clip is a measurement you cannot reproduce from what is released. At 720p the paper's ablation also shows the distilled model losing about 1.7 VQA points (67.72 against 69.37).

Memory, from Table 3: 59.5 GB at 4K inference, 179.4 GB to train. The Space runs on Hugging Face ZeroGPU and its code comments target a 48 GB budget. To get there it changes two things in its copy of wan_video_dit.py: RoPE in fp32 instead of the upstream fp64, because "at 685k tokens the fp64 cast alone allocates >16 GB transients per q/k" (lines 104-107), and the FFN applied in chunks of 65,536 tokens instead of materialising the 12 GB hidden state at once (lines 374-384). Both are mathematically harmless. It also defaults to 15 denoising steps instead of the paper's 50, and the X post's own footer reads "15 steps". Its timing constants are labelled in the code as placeholders, so I would not quote the Space's "10-20 minutes" as a measurement.

The "33 Hours" in Figure 1 is the mismatch I mentioned. Table 2's own components add up to 22.2 hours for the same model at the same resolution on the same GPU. The figure's bar chart uses Table 2's DiT numbers, so I suspect the label came from a different run. Either way the headline survives: one hour against roughly a day.

The speedup gap that is not there

The paper has a paragraph called "Gap between actual and theoretical speed": Table 1 promises 43.0× fewer MACs at 4K and Table 2 delivers only 21.4×, and at 720p it is 30.9× against 4.7×. It blames "software-hardware mismatch" and "reduced compute density from tiling".

But 43.0× and 30.9× are the attention column only. The transformer also spends 857.7 trillion MACs on projections, FFN and cross-attention at 4K, and T3 does not touch any of it. The right theoretical ratio for DiT time is the "All" column: 44,157.1 / 1,864.5 = 23.7× at 4K, and 621.4 / 111.7 = 5.6× at 720p. Measured, the paper gets 21.4× and 4.7×: 90% and 84% of the achievable ratio. Every resolution in the table lands between 84% and 90%.

Put another way, both models run at nearly the same arithmetic rate on the H20. Stock Wan does 88,314 TFLOP per 4K forward in 793.2 s, about 111 TFLOPS. T3 does 3,729 TFLOP in 37.1 s, about 100 TFLOPS. Short windowed attention calls are about 10% less efficient than one giant one. A respectable result, and a more honest framing than the paper's: T3 is close to the limit of what removing attention can buy, and the next win has to come from the FFN, which is now the biggest line item. In the calculator above, toggle "run the per-frame layers once" at 4K and watch the whole-DiT number. The FFN does not move.

What the pixels look like

The claim in the post is about detail at 100%. The X video is a 1440×1080 encode, so it cannot show native 4K pixels in a full frame, but it does switch to a 1440-pixel-wide 100% crop of the 3840-wide output. Below are the full frame and the crop, taken from that video.

A red sandstone canyon seen from above, with a dry river bed winding through scrub on the canyon floor. The full 4K frame scaled down to 1440 pixels wide.
Full frame, 3840×2176 scaled to 1440 wide. A frame from the @HuggingApps post's video, generated by the Space at 15 steps (HuggingApps on X).
A 100% crop of the canyon floor: hundreds of individual green and grey bushes on sandy ground, with the dark river channel curving down the right side.
The same video at a 100% crop: 1440×816 output pixels, one screen pixel per generated pixel, after X's re-encode (HuggingApps on X).

The bushes are individually there, with their own shapes and shadows. They are also soft: no twigs, little texture inside each shrub, and the edges have the slightly painted look of a diffusion model under-sampled at 15 steps and then compressed by X. It reads as plausible landscape at 4K, not as photographic 4K. The authors' own teaser frame, which ships in the repo as assets/teaser.jpg at 3812×2160 (so already resized from 3840), shows the same character up close:

A 100% crop of a penguin climbing out of turquoise water beside a mossy rock. The rock texture is detailed; the penguin's flipper is motion-blurred and its plumage is smooth.
A 100% crop from the authors' 4K World Vision teaser frame, 680×620 pixels. Rock texture holds up; fur and feathers are smooth (T3-Video repository, assets/teaser.jpg).

Rock and water texture hold up well. The penguin is clean but smooth, the kind of surface a 1.3B model renders as a gradient. Fine semantic detail at this scale needs capacity, and the paper says as much in its limitations: it did not try Wan2.1-14B for lack of data and compute.

Against upscalers and the other 4K attempts

The standard way to get 4K video is to generate at 720p or 1080p and run a video super-resolution model. The paper's argument against it is that super-resolution invents high-frequency detail that does not belong to the scene and flickers between frames. It is a fair argument, but the paper does not test it. It skips the comparison because UltraGen "already proved superior to cascaded" pipelines, and UltraGen, UltraWan and the UltraVideo dataset all come from overlapping author lists with T3. There is no head-to-head against a current upscaler anywhere in the paper, so the "no upscaler" framing is a design choice rather than a measured win. If I had to ship 4K video tomorrow on a budget, T3 at 720p is 259.9 s per clip on an H20 by the paper's Table 2, and a cascade from there is the obvious baseline someone should run.

The comparisons it does make, on its 120-clip 4K-VBench:

model (4K)framesVQAVTC
Wan2.1-T2V-1.3B, official8130.010.37
HunyuanVideo, official4161.920.68
UltraGen2967.430.75
T3-Video-T2V-1.3B8171.720.83
T3-Video-T2V-1.3B, LoRA8170.780.79

The headline "+4.29 VQA, +0.08 VTC" is T3 minus UltraGen, and the subtraction checks out. VQA is FineVQ, VTC is a text-consistency score from Qwen3-VL-32B. Ten professional raters preferred T3 over UltraGen 71.25% of the time on video quality and 63.42% on detail. Note the frame counts: UltraGen makes 29 frames, T3 makes 81. The "7 hours versus 1 hour" in Figure 1 is UltraGen's 29 frames against T3's 81, which flatters UltraGen if anything.

Against the efficient-attention literature, T3 sits in a particular spot. Sparse methods like Sparse VideoGen, Sliding Tile Attention and Radial Attention try to keep the model untouched and skip work at inference, and VSA trains a learned sparsity pattern. T3 picks a fixed pattern and retrains every weight to live with it. It costs more than the training-free methods and adapts less than learned routing, but the pattern is static, dense inside each window, and runs on stock FlashAttention, which is why its measured speedup tracks its MAC count so closely. The paper also swapped T3 for Swin-style shifted windows under the same training at 720p: 67.34 VQA against 69.37. And at the far end, LinGen's linear-complexity route reportedly cost around 10K H100 GPU-days of training at 512p, which puts the 64-GPU fine-tune here in perspective.

The authors also make a claim I find more interesting than the 4K result: a T3 model can be turned back into full attention and recovers in 500 iterations at 720p (69.51 VQA against the official 70.56). If that generalises, windowed attention becomes a cheap way to pretrain at high resolution and convert at the end. It rests on one table and one model, so I would treat it as a lead, not a finding.

If you have been following this site's video coverage, compare it with SANA-Video 2.0, which goes the other way and makes three of four layers linear, and Sol-Attn and FastH3, which skip attention blocks at inference. The attention field guide puts dilated and local windows in context, and FlashAttention-3 explains why the score matrix never hits memory. LoT Diffusion is the image-side cousin: a DiT whose token granularity varies by region.

Should you use it

If you want native 4K text-to-video from open weights today, this is one of very few options, and the licence is Apache-2.0 on top of Wan2.1's terms. Budget an hour of a 60 GB-class GPU per 50-step clip, or follow the Space and accept 15 steps and softer detail. You get one shape, 2176×3840 at 81 frames, and nothing else without retraining.

The idea is worth more than the checkpoint. Two windowed passes with shared weights, cycled through a few shapes, is a small, portable change you can apply to any DiT you are willing to fine-tune. If you do, run the per-frame layers once.

How I checked

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "T3-Video: 685,440 tokens per 4K clip, and the window attention that makes them affordable", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026t3video4k,
  author = {Satyajit Ghana},
  title  = {T3-Video: 685,440 tokens per 4K clip, and the window attention that makes them affordable},
  url    = {https://ai.thesatyajit.com/articles/t3-video-4k},
  year   = {2026}
}
share