~/satyajit

FreeVideo: a 66 GB video transformer on an 8 GB card, and the planner that decides where every block lives

mdjsonmcp

2026-10-06 · 20 min · video-generation · diffusion · inference-optimization · quantization · linear-attention · edge-inference · explainer

FreeVideo went open source on 2 October from FlashML-org, the group behind FreeToken. The pitch went round X as "video generation just moved onto your gaming PC": MiniMax H3 running locally on "as little as 8GB of VRAM and 16GB of RAM", with ComfyUI integration, LoRAs, and acceleration that "adapts to whatever hardware you have." Haocheng Xi, one of its five authors, framed it as MiniMax H3 plus Video DeltaNet. A Japanese repost made the same point: the result comes from combining H3 with Video DeltaNet.

That is a 33-billion-parameter audio-video transformer on a card with 8 GB. I want to know how. So I read the engine at commit 9925fe7, the Video DeltaNet paper, and the tensor headers of every checkpoint it downloads, without running anything. The short version: no single trick does it. Four separate things shrink or move the bytes, and a planner with 1,142 lines of measured special cases decides where each one goes for every request. Video DeltaNet is real and it matters, but what it buys is speed, not memory.

FlashML-org/FreeVideo@9925fe7 · snapshot 2026-10-06
tracked files
356
license
Apache-2.0
branch
main
tests
1 file
source
3.1 MB
commit date
2026-10-06
source by language
Python2.8 MB(231)JavaScript286.3 kB(22)CSS58.7 kB(11)PowerShell21.1 kB(5)Shell17.9 kB(3)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at 9925fe7 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

The FreeVideo workspace inside ComfyUI: a prompt box on the left with landscape, portrait, square and match-image presets, 1344 by 768 at 10 seconds and 24 fps, two-pass sampling ticked; on the right a generated frame of a giant ringed structure over a reflective desert at sunrise, with tiles reading 78.0 s sampling, 130.7 s request total, 29.4 GiB VRAM peak and 22.5 GiB RAM peak.
The FreeVideo creative workspace. This screenshot is from a 32 GB card (29.4 GiB VRAM budget), not an 8 GB one. The engine reports the memory peaks of every request it runs (FreeVideo README).

What MiniMax H3 is, briefly

I covered the base model when it shipped, and FastVideo's four-step distillation of it in August. The facts that matter here come from the transformer's config.json and its safetensors index, all measured:

FreeVideo's geometry.py turns a request into tokens. Frames round up to 17n+5. A 17-frame VAE clip becomes 5 latent frames. At 1344×768 every latent frame is 24×42 = 1,008 tokens. So a 10-second clip is 243 frames, 72 latent frames and 72,576 video tokens. A 14.4-second clip (345 frames) is 102 latent frames and 102,816 tokens. Both counts are hard-coded in the planner's measurement tables, which is how I know they are the ones it was tuned on.

Why it does not fit: the arithmetic

Start with the weights. 66.28 GB of BF16 against 8 GB is a factor of eight before a single activation exists.

Then the activations, at 102,816 tokens (all reasoned from the shapes above, in BF16):

TensorShapeSize
One copy of the residual stream102,816 × 5,3761.11 GB
Q, K and V for all 56 heads3 × 102,816 × 7,1684.42 GB
Feed-forward hidden (gate and up, before SwiGLU)102,816 × 28,6725.90 GB
Dense attention scores, one head, if materialized102,816²21.1 GB

No sane engine materializes the last row. Flash-style kernels stream it. But the other three are real, and they coexist with the weights of whichever block is running. The VDN paper adds the compute side: in the profiled H3 workload, softmax attention is "more than 85% of denoiser runtime" (reported).

So there are two problems. The memory problem is weights. The time problem is attention. FreeVideo attacks them with different tools.

Cut one: 26 GB of AdaLN weights become 0.22 GB of tables

The single biggest saving is not quantization. Read block 0's header from the H3 checkpoint and one tensor dwarfs the rest: adaln_proj.linear.weight, shape 96,768 × 2,688. That is 260.2M parameters per block, more than the attention (154.1M) or the feed-forward (231.2M). Across 50 blocks it is 26.02 GB in BF16 (reasoned), about 39% of the transformer. This is the "about 13B of the 33B" of modality-specific AdaLN branches MiniMax's card mentions.

Here is the observation that kills it. The input dimension is 2,688, which is time_embed_dim. AdaLN modulation in H3 depends on the timestep embedding and nothing else. The output is 96,768 = 3 modalities × 6 vectors × 5,376 channels: a shift, scale and gate for attention and for the feed-forward, per modality. If the sampler always visits the same timesteps, the modulations are constants.

VDN-H3 is an eight-step distilled sampler. Its timesteps are fixed. So adaln.py evaluates the original projection once per schedule row, stores the result, and replaces the module with CachedModulation, which indexes a buffer by step. The docstring credits the idea to NVIDIA's Sol-H3 AdaLN precompute, and the VDN paper's own serving stack does the same thing. A ScheduleCursor checks that each call's timestep equals the stored one and raises if a request changes the schedule.

The prepared tables ship in OpenVDN/vdn-minimax-h3-edge: 50 files of 4,456,040 bytes each, 0.22 GB in total (measured). The model card says the original projections are omitted, "saving approximately 26.02 GB per installation", which matches my arithmetic to the hundredth. What is left of the BF16 transformer is 40.26 GB.

The cost is flexibility, and the card says so. A schedule other than the published ones, or a LoRA that changes the modulation, needs the original projections from a pinned older revision. Ordinary attention and feed-forward LoRAs keep the tables. The four quality levels added in v0.2.0 (Light is 8 steps in two passes, Medium 12, High 16, Max 20) each need their own rows; policy.adaln_extra_bytes charges them to VRAM at 3 × 6 × 5,376 × 2 bytes per modality timestep per layer.

Cut two: FP8 for the seven wide matrices

The second cut is ordinary. FreeVideo's prepared block file is 432,433,052 bytes. Its header lists exactly seven FP8 E4M3 matrices: the original Q, K, V and output projections, a second output projection for the linear branch, and the two feed-forward matrices. Together they are 423.9 MB. Everything else in the block (norms, gates, the linear branch's short convolutions and decay projections) stays in BF16 and adds 8.5 MB. All 50 blocks plus a 0.93 GB root file come to 22.55 GB (measured), which lines up with the "about 22.9 GB" the model card says FreeVideo downloads.

bytes on disk, decimal GB · bar width out of 70 GB
the denoiser
H3 transformer, BF16 as released66.28 GB measured
33.14B parameters · 50 blocks + 2 refiner blocks
… without the AdaLN projections40.26 GB reasoned
50 × 96768×2688 weights replaced by 8-step tables (0.22 GB)
… with the wide matrices in FP822.55 GB measured
50 × 432.4 MB blocks + 0.93 GB root · VDN branch included
what runs around it
Qwen3-VL-32B text encoder, BF1666.71 GB measured
MiniMaxAI/MiniMax-H3 text_encoder shards
… NVFP4 AWQ, as FreeVideo installs it15.69 GB measured
runs in its own process, exits before the transformer loads
Video VAE, BF1610.42 GB measured
decoder is 36 transformer blocks; streamed below a 20 GiB budget
Even after both cuts the transformer is 22.55 GB, nearly three times an 8 GiB card. Nothing here makes it fit. It makes it small enough to stream: with 16 GiB of system RAM even the host cannot hold all of it, and the project's own grid marks that configuration as re-reading part of the stack from disk on every step.

The interesting part is that FP8 is not one code path. docs/execution-planning.md and fp8_gemm.py pick per architecture:

The model card is careful about what this proves. The per-tensor package was checked bit-identical against the full prepared model on a 362-frame request. The rowwise package "adds no new physical Hopper, Ada or Windows benchmark." That is an honest sentence and I wish more model cards had one.

Cut three: the text encoder never meets the transformer

The Qwen3-VL-32B encoder is 66.71 GB in BF16. FreeVideo installs a community NVFP4 AWQ conversion of it, 15.69 GB. That still does not fit, and it does not need to. The encoder runs in a separate process that exits before the transformer loads. Its only lasting requirement is host memory: the planner refuses any machine with less than 3.6 GiB of RAM budget for it, a figure measured on an H200 after the encoder's checkpoint pages were released. Cached or pre-encoded prompts skip it entirely.

The rest is placement

After the three cuts, 22.55 GB of transformer still has to pass through an 8 GB card for every one of the eight denoising steps. This is where "adapts its acceleration to whatever hardware you have" lives. It turns out to be a measured policy, not a learned one.

Before every request the planner reads free VRAM, available host memory (including cgroup limits and, on Windows, commit headroom), the GPU's compute capability, and which attention kernels pass an on-device probe. It turns the request into a token count. Then it decides five things.

1. The head group. Attention runs on 4, 8 or 16 of the 56 heads at a time, and the QKV projections are sliced to match (head_chunk.py). A slice is tokens × heads × 128 × 2 bytes. Four heads at 102,816 tokens is 105 MB per tensor instead of the 4.42 GB above. policy.py holds a table of what each group measured beside the resident weights: at 72,576 tokens, 7.69 GiB for four heads, 8.03 for eight, 10.05 for sixteen. The planner picks the widest group that fits. The comments say why it bothers: a wider group is worth 13% to 51% where it fits.

2. How many blocks stay on the GPU. ⌊(VRAM budget − activations) / 432.5 MB⌋, capped at 44, then pulled down to 8 whenever host RAM can hold the rest. The cap and the floor both come from measurements in the source. At a 24 GiB cap, 8 resident blocks ran at 9.20 s/step and 29 ran at 9.24. At 32 GiB, 8 blocks ran at 9.15 and 48 at 9.28. Streaming the remainder moved 19.3 GiB in 0.44 s over a host link measured at 47.7 to 54.1 GB/s. The transfer is not the bottleneck; the resident copies just take memory away from everything else.

3. Where the other blocks wait. Non-resident blocks sit in pinned host memory, up to 22 GB, and are copied to the GPU through one or two transfer slots on every step (offload.py). A second slot, for prefetch, only appears at a 14 GiB budget or more, with lower thresholds on some Windows and Ampere setups. If host RAM cannot hold them either, a bounded subset stays pinned and the rest is re-read from disk on every step.

4. Activation staging. Below a 10 GiB budget, if the activation path does not fit, the residual stream and attention outputs move to pinned host buffers. The code treats this as a last resort. On an 8 GiB card at 243 frames, holding no weights at all and keeping activations on the card ran at 13.49 s/step against 34.76 for staging, the 2.6× the comments keep citing.

5. The VAE decoder. Below a 20 GiB budget, ⌊(budget − 3.25 GiB) / 268.6 MB⌋ of its 36 blocks stay resident and the rest stream. Temporal clips decode one at a time, and finished frames leave the GPU before the next clip starts (decode_stream.py).

The widget below applies those rules. It is a sketch, and it says where it simplifies. The fitted host-headroom constant is mine, not FreeVideo's. With it, the sketch reproduces every cell of the project's published blocks-in-VRAM grid at the sizes it offers.

1344×768 · 243 frames · 72576 video tokenssketch of the documented rules · Linux
VRAM
RAM
clip
GPU budget 7.50 GiBhead group 4 of 56 · small-card band
■ activations ≈ 7.69 GiB ■ 0 resident blocks = 0.00 GiB
where the 50 FP8 transformer blocks liveRAM budget 15.00 GiB
■ 0 in VRAM ■ 31 in host RAM ■ 19 re-read from disk
over PCIe per step
21.6 GB
from disk per step
8.2 GB
VAE decoder blocks on GPU
16 of 36
H200 grid, s/step
13.80

Residency is capped at 44 and pulled down to 8 whenever host RAM can hold the rest: the project measured 8 blocks running 0.4% and 1.3% faster per step than 29 and 48. Blocks the host cannot retain either stay on the GPU or come off the disk every step.

A simplified re-implementation of the rules in FreeVideo's docs/execution-planning.md and policy.py, with free VRAM taken as the full card and a fitted 2.5 GiB of host working memory. It reproduces the published blocks-in-VRAM grid for the 10 s clip at these sizes; the s/step column is the project's own H200 measurement with memory capped, so it shows the cost of placement, not a consumer card's speed.

The point the widget makes is about the bottom-left corner. An 8 GiB card at 10 seconds holds no transformer blocks at all. It keeps only the activation path, about 7.7 GiB, and every one of the 50 blocks crosses PCIe on every step, 21.6 GB of it (reasoned). With 16 GiB of RAM the host cannot keep all of them either, so part of the stack is also re-read from disk each step. The "8 GB + 16 GB" headline describes that configuration. Unlike most low-VRAM claims, including WanGP's, it names the RAM, and the planner refuses the request outright rather than spill when even that is not there.

A heat-map grid with VRAM from 8 to 32 GiB on the x axis and system memory from 8 to 32 GiB on the y axis, each cell giving the number of transformer blocks held in VRAM. Small cards hold 0 to 3 blocks; large cards with plenty of RAM settle at 8; large cards with little RAM hold up to 39. Dots mark cells that re-read blocks from disk every step, covering the low-RAM, low-VRAM corner.
Transformer blocks kept in VRAM for every combination of 8 to 32 GiB of VRAM and host memory, at 1344×768 and 243 frames. Note the plateau at 8 once RAM is plentiful, and the dots where blocks come off the disk every step (FreeVideo, docs/execution-planning.md capacity grid).
A matching heat-map of seconds per sampling step. The 8 GiB VRAM column ranges from 12.6 to 15.1 seconds, worst at 8 GiB of RAM; from about 15 GiB of VRAM with 16 or more GiB of RAM, almost every cell sits near 8.9 to 9.1 seconds.
Seconds per sampling step for the same grid, measured on one H200 with VRAM capped by an MPS client limit and RAM by a cgroup. Compute and bandwidth are an H200's throughout, so this isolates the cost of placement, not a consumer card's speed (FreeVideo, docs/execution-planning.md capacity grid).

Read the two grids together and the cost of fitting is smaller than I expected. On the H200, going from 32 GiB of VRAM to 8 costs 8.93 → 12.7 s/step with plenty of RAM, and 15.1 with 8 GiB of RAM, where the disk is in the loop (reported). That is 1.4× to 1.7× slower, not 10×. The catch is in the caption: every cell has an H200's compute and host link behind it. A consumer card computes more slowly, which hides transfers better, and has a narrower link, which hides them worse. The RTX 4060 Ti, for one, has an eight-lane PCIe 4.0 link, 16 GB/s nominal, so streaming all 50 blocks takes at least 1.35 s per step (reasoned) where the H200 took a fraction of that.

Where Video DeltaNet comes in

All of the above is about memory. Video DeltaNet (paper, Xi et al., UC Berkeley, Impossible and UT Austin, September 2026) is about the 85% of runtime that attention eats. FreeVideo does not train anything; it runs OpenVDN's released VDN-H3 checkpoint, which adds a linear-attention branch and two LoRA adapters to the frozen H3 backbone.

The split is by temporal distance. Softmax attention stays exact, but only within a window: each query chunk of five latent frames (one VAE chunk) sees itself and its two neighbours, a 15-frame window. Two boundary anchors add global connectivity: every frame sees the first and last latent frames, and those two frames see everything. Text and audio interactions stay fully softmax. Everything else, the distant video context, goes to a linear branch.

Two attention-pattern grids side by side for a toy sequence of one text token, twenty video tokens in ten two-frame chunks, and one audio token. The softmax grid has a banded diagonal of local video windows, full first and last video rows and columns as boundary anchors, and full text and audio rows and columns. The linear grid fills the remaining off-band video region in blue, with arrows showing a forward scan from the left and a reverse scan from the right.
How VDN divides H3's attention. Softmax keeps local windows, boundary anchors, text and audio; linear attention covers distant video with a forward and a reverse scan (Video DeltaNet paper, Figure 2).

The linear branch is a delta-rule memory in the family of Gated DeltaNet, run as two scans, one forward and one reverse, so a frame reads a summary of the past before its window and of the future after it. The prompt is folded in too: a text state STS_T seeds each scan with ST/2S_T/2, so the two readouts together count it once.

The part that is new is the update. A language-model delta rule writes one token at a time. A video frame arrives as 1,008 tokens at once, with no natural order. Video Delta Attention makes the frame's new state the solution of one joint least-squares problem. With Sˉt\bar S_t the decayed state, kt,u,vt,uk_{t,u}, v_{t,u} the keys and values of the frame's tokens and βt,u\beta_{t,u} their write strengths:

St=arg⁡min⁡S12∥S−Sˉt∥F2+12∑uβt,u∥Skt,u−vt,u∥22S_t = \arg\min_S \tfrac{1}{2}\lVert S - \bar S_t \rVert_F^2 + \tfrac{1}{2}\sum_u \beta_{t,u}\lVert S k_{t,u} - v_{t,u}\rVert_2^2

which has the closed form

St=(Sˉt+Bt)(I+At)−1,At=∑uβt,ukt,ukt,u⊤,Bt=∑uβt,uvt,ukt,u⊤.S_t = (\bar S_t + B_t)(I + A_t)^{-1},\quad A_t = \sum_u \beta_{t,u} k_{t,u} k_{t,u}^\top,\quad B_t = \sum_u \beta_{t,u} v_{t,u} k_{t,u}^\top .

The inverse is only dk×dkd_k \times d_k, 128 × 128, so it is cheap. It is what lets two tokens with overlapping keys account for each other instead of writing conflicting corrections. The paper also proves that the carried-over part of the state cannot grow from one frame to the next, because the norm of Diag(αt)(I+At)−1\mathrm{Diag}(\alpha_t)(I + A_t)^{-1} is at most 1. The linear-attention roundup has more on why the shape of that update matters.

Two block diagrams. Left, the hybrid attention layer: the input feeds a softmax branch and a linear branch; each is multiplied by its own sigmoid gate, the linear branch passes through RMSNorm, each has its own output projection, and the two are summed. Right, the linear branch: Q, K and V from linear projections, a short convolution on K and V, SiLU, L2 normalization on Q and K, a decay alpha and a write strength beta, all feeding Video Delta Attention with a forward and reverse scan, seeded by half the text state.
The hybrid layer and its linear branch. Each branch has its own gate and output projection; the feed-forward sublayer is unchanged (Video DeltaNet paper, Figure 3).

You can see this directly in FreeVideo's block file. Besides the seven FP8 matrices, it holds linear_attention.alpha, beta_proj, output_gate, the 5×5 spatial and 5-tap temporal short_conv filters on K and V, a softmax_gate, and the second output projection to_out_linear. The branch adds about 43M parameters per block (reasoned).

Here is the honest accounting for memory. The linear state per head is 128 × 128, which is tiny. But flash-style kernels already kept dense attention's memory linear in sequence length, so VDN does not shrink the activation footprint in any way that matters. It makes the call cheaper. The paper's numbers, all reported:

Top: a first-last-frame-to-video comparison at 14.3 seconds and 768p, a man holding a white cat above a canyon city, split diagonally between dense H3 at 50 steps and VDN-H3 at 8 steps, with six more matched frames. Bottom left: attention latency against video duration, dense H3 rising from about 6 to 31 seconds while VDN stays below 10, with speedups of 1.77x, 2.57x and 3.38x. Bottom right: denoising time on B200 falling from 799.6 s dense to 307.9 s with VDA, 49.3 s with eight steps and 6.70 s on eight GPUs.
VDN-H3 against dense H3: matched frames, attention latency by clip length, and the cumulative denoising speedup on B200 (Video DeltaNet paper, Figure 1).

Those quality numbers are for VDN-H3 on datacenter GPUs. FreeVideo slices heads, stages buffers and swaps kernels, and its own docs say an out-of-memory retry "can change the floating-point reduction order." FreeVideo publishes no quality metrics of its own for its plans or for its four levels. The gallery shows the levels side by side. That is evidence you can look at, not a measurement.

Which attention kernel runs

FreeVideo probes attention backends on the device and takes the first that passes: SageAttention 2, PyTorch flash attention, cuDNN, FlashAttention 2, FlashAttention 4 (attention.py). Probe results are cached against a hash of the GPU, driver, packages and compute source, so an upgrade re-probes. SageAttention quantizes Q and K internally, so the default consumer path is not bit-identical to the paper's kernels. Each request's .request.json records which backend ran, but nothing in the repository measures the quality difference.

What it costs on real hardware

The README's end-to-end table is for Windows, 1344×768, 10 seconds, two-pass sampling. Two-pass means eight steps at half resolution, a latent upscaler, then three steps at full size. All of these are reported:

GPUVRAM + RAMTime
GeForce RTX 509032 GB + 64 GB122 s
GeForce RTX 5060 Ti16 GB + 32 GB486 s
GeForce RTX 4060 Ti16 GB + 32 GB558 s
GeForce RTX 4060 Ti (community report)8 GB + 64 GB603 s

The 8 GB row is a community report with 64 GB of RAM, not the 16 GB the headline quotes, so the README itself does not show the exact 8 + 16 configuration end to end. The H200 grid covers it, at 13.8 s per full-resolution step with compute that is not a consumer card's. The 8 GB 4060 Ti being only 8% slower than the 16 GB one is the residency result again. Once the stack is streaming, extra VRAM buys little.

On Apple silicon, the Mac guide reports an M5 with 24 GB of unified memory at about 42 minutes for the same 10-second clip, using ConvRot int8 weights on the M5's tensor units.

What I would watch

So FreeVideo does not fit MiniMax H3 into 8 GB. It deletes a third of the model by noticing that the third depends only on the timestep, stores the rest at a byte per weight, and rents the card one 432 MB block at a time on a plan it has already measured. Video DeltaNet makes each rented step about three times cheaper. The same placement-over-shrinking pattern runs through FreeToken and stable-diffusion.cpp; SANA-Video 2.0 and MAGI-2 attack the attention bill from the training side, and the diffusion transformer doc covers what all of them rest on.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "FreeVideo: a 66 GB video transformer on an 8 GB card, and the planner that decides where every block lives", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026freevideominimaxh3,
  author = {Satyajit Ghana},
  title  = {FreeVideo: a 66 GB video transformer on an 8 GB card, and the planner that decides where every block lives},
  url    = {https://ai.thesatyajit.com/articles/freevideo-minimax-h3},
  year   = {2026}
}
share