~/satyajit

Kandinsky 6.0 Video: two token streams on one clock, and where the lip-sync comes from

mdjsonmcp

2026-10-06 · 27 min · video-generation · audio · diffusion · flow-matching · reinforcement-learning · distillation

Why read this

Notabletop 60%

Explains joint audio-video denoising from the code (230:1 tokens, RoPE scaled 0.144) and shows the 47% WER cut is graded by its own reward model.

  • Original analysis
  • Runs on a consumer GPU
  • Concrete numbers to act on

Image & video generationMITPractitioner model

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
2 of 3: Open code or weights with real limits
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
1 of 3: Relevant for months
Does it affect many?
1 of 3: A specialist community
Only here?
2 of 3: A teardown or measurement few others did

Score 59 of 100, ranked 234 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Kandinsky Lab released Kandinsky 6.0 Video this week: two diffusion models that write a 5-second clip and its soundtrack in one pass. Lite is advertised at 3B parameters, Pro at 29B. Both take text, or text plus a first frame, and return SD video with 44 kHz audio, speech and lip movement included. A separate super-resolution model takes the picture to 1920 x 1080. Weights, code and a diffusers integration are MIT-licensed. The paper is unusually complete. It gives data filters, per-stage learning rates, reward weights and timings on seven GPUs.

This piece explains how a model draws a picture and its sound together, from the token up. Then it checks the release against what can be read without a GPU. I read the paper, the kandinskylab/kandinsky-6 code at commit ea44d1b, the diffusers configs, and the safetensors headers of the three ungated transformers, over HTTP range requests. The non-distilled Pro checkpoint is gated behind a click-through. Its parameter count comes from the Hub API, which reads the same header. I ran nothing on a GPU. Every number below is labelled measured (I read it from a file), reported (the paper's figure) or reasoned (my arithmetic on the other two).

kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers@0882a24 · snapshot 2026-10-06
announced
29B
measured
30,136,326,248
parameters
30.14B
repo size
223.47 GB
task
image-to-video
library
diffusers
license
mit
safetensors
10 shards
largest file
69.73 GB
files
40
downloads
173
likes
15
gated
auto
parameters by dtype
BF1625.41BF324.73B
videovideo-generationtext-to-videoimage-to-videoaudio-generationkandinskydiffusers

Transformer only: 30,136,326,248 parameters, 4,730,755,072 of them stored in F32 (measured, Hub API), which is why the file is 69.7 GB. The distilled Pro is 30,139,423,760 parameters, 60.3 GB, almost all BF16 (measured from its header). Gated click-through; MIT licence.

repo last modified 2026-10-06

kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers@6510114 · snapshot 2026-10-06
announced
3B
measured
3,176,632,424
parameters
3.18B
repo size
35.25 GB
task
image-to-video
library
diffusers
license
mit
safetensors
10 shards
largest file
7.48 GB
files
40
downloads
53
likes
9
parameters by dtype
BF162.62BF32561.5M
videovideo-generationtext-to-videoimage-to-videoaudio-generationkandinskydiffusers

Transformer only: 3,176,632,424 parameters (measured from the header). The repo also ships the Qwen2.5-VL-7B and CLIP text encoders, the Hunyuan video VAE, and the MMAudio VAE and vocoder. Ungated; MIT licence.

repo last modified 2026-10-06

kandinskylab/kandinsky-6@ea44d1b · snapshot 2026-10-06
tracked files
193
license
MIT
branch
main
tests
2 files
source
631.3 kB
commit date
2026-10-06
source by language
Python605.8 kB(122)CUDA10.3 kB(1)C6.7 kB(2)JavaScript5.4 kB(2)Shell2.0 kB(1)C++1.0 kB(1)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at ea44d1b — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

Four frames in a two-by-two grid: a brass deep-sea diver among fish in a flooded cathedral; a lion in a green coat on a subway car between passengers on their phones; a young woman blowing a pink bubblegum bubble in an autumn park; a reptile's head in a rainforest between two gloved hands.
Four frames from the model card's promo reel. The reel is the publisher's selection; it has no audio on the page and says nothing about how often a prompt comes out like this (Kandinsky 6.0 Video model card).

Two ways to put sound on a picture

The obvious way is a cascade. Generate the video, then hand it to a video-to-audio model that listens to the pixels and writes a soundtrack. MMAudio, whose autoencoder Kandinsky 6.0 reuses, is that second stage. A cascade is modular, and each half can be trained on its own data. It also gets one thing structurally wrong. Speech has to drive the mouth, and a cascade can only drive it the other way. By the time the audio model runs, the lips are frozen pixels. The audio can only chase them.

The joint way denoises both at once. Every step, the video sees where the audio is going and the audio sees where the video is going, so a word and the jaw that shapes it are one decision. There are two designs. One puts everything into a single token sequence with modality-specific weights, which is what MAGI-2 does. The other keeps two transformers, one per modality, and has them talk through cross-attention in every block. Kandinsky 6.0 is the second kind, like Ovi and LTX-2. I took apart LTX-2's version of that block in the restore and relight LoRAs piece. MiniMax's H3, which FastH3 and FreeVideo work on, is another large joint audio-video transformer.

One block, two streams

A block diagram with two rows. The top row, on the video embedding: Video Self Attention, T2V Cross Attention, V2A Cross Attention, Feed Forward. The bottom row, on the audio embedding: Audio Self Attention, T2A Cross Attention, A2V Cross Attention, Feed Forward. Text feeds both text cross-attentions as keys and values; video feeds the A2V block as keys and values and audio feeds the V2A block.
One Kandinsky 6.0 Video block. Each stream runs self-attention, text cross-attention, a cross-modal attention and an FFN; in V2A, video queries attend to audio keys and values, and A2V is the reverse (paper, Figure 1).

Read in the code (FusedTransformerDecoderBlock in dit.py), one block does this, in order:

  1. Self-attention, per stream. Video tokens attend to video tokens with a 3D RoPE over (time, height, width). Audio tokens attend to audio tokens with a 1D RoPE over time.
  2. Text cross-attention, per stream. Each stream has its own small text transformer over the Qwen2.5-VL embedding, so the video and the audio read the same prompt through different weights.
  3. Cross-modal attention, both ways. Video queries attend to audio keys and values, and audio queries attend to video. The keys and values are taken from the other stream's state after self-attention but before its text cross-attention, and each result is added back through a learned gate (measured, dit.py).
  4. Feed-forward, per stream.

Each stream gets its own timestep embedding. During joint training the noise level is drawn independently for video and audio, "so that the two modalities are noised to different levels" (reported). This is what lets one network also do audio-to-video and video-to-audio: hold one stream at zero noise and denoise the other.

The two streams are deliberately unequal. The paper gives the split in round numbers. Here it is from the headers:

Lite (paper)Lite (header)Pro (paper)Pro distill (header)
video stream2B1,909,391,36019B18,375,260,160
audio stream0.6B543,657,9845B4,909,455,360
cross-attention0.4B595,152,8965B5,664,921,600
text blocks and embeddingsnot counted128,430,184not counted1,189,786,640
total3B3,176,632,42429B30,139,423,760

The header columns are measured. I bucketed every tensor by its prefix. "Cross-attention" means the va_ and av_ attention and modulation tensors in every block. The total row of the paper is the sum of its own three lines; it leaves out the per-stream text transformers, which is most of the 1.1B gap on Pro (reasoned). The cross-attention is 18.7% of Lite and 18.8% of Pro (reasoned from the measured counts). Lite has 32 blocks with a 1,792-wide video stream and an 896-wide audio stream. Pro has 60 blocks at 4,096 and 2,048 (measured, transformer/config.json).

Counting one clip

Now the token budget, because it explains most of the design. Video goes through the Hunyuan VAE, which compresses 8x in space and 4x in time into 16 channels. The transformer then patches 2 x 2. One video token is a 16 x 16 pixel square of one latent frame. The default clip is 121 frames at 24 fps, which the causal VAE turns into 31 latent frames: the first frame alone, then four per latent. At 864 x 480 that is 54 x 30 = 1,620 tokens per latent frame, and 50,220 video tokens per clip (reasoned from the configs and pipeline.py).

Audio goes through MMAudio's 44.1 kHz autoencoder, which turns 1,024 samples into one 40-channel latent frame. The DiT does not patch it: each latent frame is one token. The pipeline sizes the audio to the video as ceil(121 / 24 x 44,100 / 1,024) = 218 tokens (measured, prepare_latents.py). That is one audio token per 23.2 ms against one video latent frame per 166.7 ms, about 7.2 audio tokens under each latent frame (reasoned).

So a clip is 50,220 video tokens next to 218 audio tokens, about 230 to 1. The audio stream holds 16% of Pro's weights and sees 0.4% of the tokens. Counting each stream's weights against its own tokens, audio is about 0.12% of the linear-layer work per step (reasoned). The audio stream is cheap to make wide. The cross-attention is where the two sizes meet: 50,220 x 218 query-key pairs each way, against 2.5 billion pairs in video self-attention.

resolution (width x height)

SD landscape, the release default

model

video stream 18.38 G · audio stream 4.91 G · cross-attention 5.66 G parameters

video: 31 latent frames x 1,620 tokensaudio: 218 tokens, one per 1,024 samples (23.2 ms)0 s1 s2 s3 s4 s5 s
video tokens50,220audio tokens218 (5.062 s)video : audio230 : 1audio share of stream FLOPs0.12%
attention pairs, video self2.52 Gattention pairs, audio self47,524attention pairs, each cross10.9 Maudio tokens under frame k8 (a = 102-109)
RoPE positions video latent frame 15 rotates its queries at temporal position 15. The audio tokens under it sit at positions 14.69-15.70 (index x 0.144). A time-exact scale would be 0.1393, which would put them at 14.21-15.18.
Token counts follow the release's own code and configs: 16 x 16 px per video token, four pixel frames per latent frame after the first, one audio token per 1,024 samples at 44.1 kHz. The release generates 31 latent frames; shorter clips are shown for the arithmetic, not as a supported setting. Stream FLOPs are 2 x parameters x tokens for each stream's own weights, ignoring attention and cross-attention (reasoned). Parameters measured from the safetensors headers.

Pick a video latent frame to see which audio tokens play under it, and where RoPE puts them.

How does a video token know which audio tokens are simultaneous with it? Through RoPE. The model passes rotary positions into the cross-modal attention too (ca_rope: true). Video queries carry their 3D position; audio keys carry their 1D one. The paper says the video frame index is normalized to 24 fps, and the audio RoPE's frequencies are all multiplied by audio_freqs_scaling = 0.144 (measured, transformer/config.json and rope.py). That squeezes 218 audio positions into roughly the range of 31 video positions: the last audio token lands at 217 x 0.144 = 31.2 (reasoned).

Two details in the code make this a hint, not a lock (reasoned, from rope.py):

Neither is a bug. A cross-attention can learn to read a consistent, slightly off signal, and the sync results below suggest this one does. But the alignment is learned, not built in.

The first-frame mode reuses the same machinery. In image-to-audio-video the reference frame's latent is appended along time, flagged by a mask channel and a learned "token role" embedding. The paper's ablation found the RoPE position does the binding. Give the reference a later temporal position and the model uses it as that later frame. The mask separates conditioning from noise. The role embedding is "partially redundant" (reported). The release code gives the appended frame position 0 (measured, pipeline.py).

Where 44 kHz comes from

"44 kHz audio" is the sample rate of MMAudio's autoencoder: 44,100 Hz. That stack is a mel-spectrogram VAE plus a BigVGAN vocoder. The vocoder's upsampling rates multiply to 8 x 4 x 2 x 2 x 2 x 2 = 512 samples per mel frame, over 128 mel bins (measured, vocoder/config.json). The VAE halves the frame rate again to reach 1,024 samples per latent (reasoned from the two configs). The DiT never sees a waveform. It denoises 218 x 40 numbers; the audio VAE decodes them to a mel spectrogram and the vocoder turns that into sound. Audio fidelity is therefore bounded by a frozen, borrowed codec. The paper's contribution is the transformer that writes into its latent space.

Training: two towers, then a merger

A flow chart. Video tower training: start from the pretrained Kandinsky 5.0 Video model, then T2V pretraining with T2AV captions. Audio tower training: T2A pretraining with T2A captions, then T2A pretraining with T2AV captions. Both feed Fused Audio-Video training: T2AV pretraining, T2AV and I2AV pretraining, model soup SFT, RL-based post-training, distillation.
Training pipeline: the video tower starts from Kandinsky 5.0, the audio tower from random weights, and the two are fused for joint training, SFT, RL and distillation (paper, Figure 5).

The video stream starts from a Kandinsky 5.0 checkpoint. The audio stream starts from random weights and learns sound alone first, on 40 million audio tracks of 5 to 20 seconds (reported). Then the newly initialised cross-attention layers connect them. Lite trains jointly for 57.5k steps and Pro for 49k, on 7 million audio-video segments. The last 4k to 5k steps put 25% of samples in image-conditioned mode (reported). A batch can also hold audio only, video only or images only (reported).

There is no lip-sync module

Nothing in the architecture is about faces. Lip-sync is pressed in from four directions, all reported:

Ten frames of a young girl in a white medical coat with a stethoscope in a bright hospital room, her mouth in different positions as she speaks, above a green audio waveform whose bursts line up with the frames where her mouth is open.
A Pro generation for the line 'When I grow up, I want to be a doctor!', frames sampled across 5 seconds above the generated waveform. One sample chosen by the authors (paper, Figure 28).

The RL stage, and what the 47% measures

A diagram of the reward phase. A video prompt and audio prompt produce N videos with different seeds, scored by a table of eight reward models: HPS v3 (video, 0-15), VideoAlign (text+video, -2 to 2), Audiobox Aesthetics (audio, 0-1), CLAP (text+audio, -1 to 1), AV-Desync via Synchformer (video+audio, 0-1), AV-Align (video+audio, 0-1), LatentSync-gated via StableSyncNet (audio, 0-1) and Whisper ASR (text+audio, 0-1). Scores are normalized, routed into clamped video and audio advantages, broadcast to tokens, and mapped to per-token rewards between 0 and 1.
The reward phase: eight reward models score each group of rollouts; the scores become separate video and audio advantages, clamped to plus or minus 5 and spread over each stream's tokens (paper, Figure 7).

The RL recipe is OmniNFT, originally built for LTX-2 and adapted here (reported). For each prompt Pro generates a group of K = 10 clips that differ only in seed (8 for Lite). The sampler is deterministic, DPM2 at 20 steps, so the seed is the only source of variety. Each of eight rewards is centred within the group and scaled by the batch's standard deviation, like GRPO. The weighted advantages then split by modality. Video rewards (HPSv3, VideoAlign) go to video tokens and audio rewards (Audiobox Aesthetics, CLAP, Whisper ASR) to audio tokens. The three sync rewards go to both. Each token gets a reward between 0 and 1. The update is not a policy gradient over log-probabilities, which a flow model does not have cheaply. It re-noises the finished clip, asks the current and previous LoRA to denoise it, and pulls the prediction toward the rollout in proportion to that reward:

ℓi=ri MSE(x0cur,x0)+(1−ri) MSE(2x0old−x0cur,x0)\ell_i = r_i \,\mathrm{MSE}(x_0^{cur}, x_0) + (1 - r_i)\,\mathrm{MSE}(2x_0^{old} - x_0^{cur}, x_0)

where rir_i is token ii's reward, x0x_0 the rollout, and x0curx_0^{cur}, x0oldx_0^{old} the current and previous predictions (simplified: the paper also normalizes each term). A high reward pulls the current model toward the rollout. A low reward pulls its mirror image toward it, which pushes the model away. A small KL term (β=10−4\beta = 10^{-4}) keeps it near the SFT model. Only a LoRA trains: rank 16, 199M parameters on Pro (0.69%) and 47M on Lite (1.57%) (reported). There is also a hand-tuned schedule that shrinks gradients flowing into the audio stream through V2A attention in some blocks, so video rewards do not wreck the audio.

Now the 47%. The paper's sentence in full: Pro's word error rate on 64 held-out speech prompts, three seeds each, drops "from 0.235 for the SFT model to 0.124, a relative reduction of 47%" (reported, p under 0.001). Lite drops from 0.179 to 0.127, 29% (reported). The arithmetic checks.

The catch is in the next line of the PDF: "WER is computed with Whisper-large-v3, which is also one of the reward models." Speech intelligibility is rewarded twice in the stack. Once directly, as Whisper ASR at weight 0.5. Once inside the lip-sync term, which multiplies the SyncNet score by 0.5+0.5 sASR0.5 + 0.5\,s_{ASR}, where sASRs_{ASR} is one minus the WER. Grading an RL policy with the model it was trained against measures how well it learned to please that model. The paper knows this and runs two other checks:

Stacked bars for six criteria comparing Kandinsky 6.0 Video Pro after RL against the SFT model on a speech-focused text-to-audio-video benchmark. Prompt following video: RL 0.31, tie 0.69. Prompt following audio: RL 0.08, tie 0.92. Visual quality: RL 0.58, tie 0.25, SFT 0.17. Audio quality: RL 0.08, tie 0.84, SFT 0.08. Audio-video sync: RL 0.41, tie 0.42, SFT 0.17. Motion realism: RL 0.50, tie 0.33, SFT 0.17.
Human side-by-side, RL vs SFT, speech-focused prompts. The largest gains are visual quality and motion; audio quality and audio prompt following are mostly ties, and the chart has no speech-quality bar (paper, Figure 21a).

The human side-by-side of RL against SFT is the better evidence that RL helped, and it did. Visual quality goes 0.58 to 0.17 for RL, and audio-video sync 0.41 to 0.17 (reported, read off the chart). The text says speech quality "improves significantly" on this benchmark, but the published chart has no speech-quality bar. Audio prompt following is 92% ties. So the RL stage clearly made clips look better and sync better. How much it improved speech rests on a metric graded by its own reward.

Ten steps instead of a hundred

The full Pro samples with 50 Euler steps at guidance 5.0 (measured, pro.yaml). With classifier-free guidance that is 100 transformer forwards. With MagCache, Pro SD evaluates 49 of those 100 (reported). The distilled checkpoint, which runs the public demo, uses 10 forwards at guidance 1.0 (measured, pro-distill.yaml). It gets there in two stages. π-Flow trains a student whose output head predicts ten velocities at once, a small policy that integrates a segment without calling the network again. The headers show it: the distilled output layers are 640 wide for video and 400 for audio, ten times the 64 and 40 of the base heads (measured). Then 300 steps of adversarial fine-tuning (Sim-LADD) on the student's own rollouts restore the texture that trajectory compression smooths away.

In a human side-by-side on image-to-audio-video, the distilled Pro is preferred 51% to 49% over the full model, and no criterion differs significantly (reported). The demo is not a downgraded model, on the publisher's evidence.

From 480p to Full HD

A pipeline diagram titled SR Tiled Inference with Latent Upscaler. LQ video pixels go through a KVAE encoder to LQ latents, split into a latent tile grid. For each tile: a Latent Upscaler x2 or x4, partial re-noising of the upscaled latent, then an SR DiT with Euler steps in an iterative denoise loop, a KVAE decoder, and an HQ tile. Tiles are merged with a Hanning blend into the HQ video.
Super-resolution: encode once, upscale each latent tile with a convolutional Latent Upscaler, restore it with a text-free SR-DiT, decode, and blend the tiles with Hann windows. The released model replaces the Euler loop with two network evaluations (paper, Figure 13).

The base model stops at 864 x 480. Full HD is a separate, text-free pipeline in a different latent space: K-VAE, 64 channels, 16x spatial compression. The clip is pre-scaled 1.125x in pixels to 976 x 544, then encoded once and tiled. A convolutional Latent Upscaler doubles each tile in latent space. The tile is partially re-noised, and a 1.41B-parameter SR-DiT, initialised from Kandinsky 5.0 Video Lite, restores detail. The output is 1952 x 1088, resized to 1920 x 1080 (reported). The released SR-DiT is distilled to two network calls per tile; its checkpoint is 1,414,263,936 parameters (measured, Hub API). On an H100, Full HD from 864 x 480 takes 9 tiles and 39.6 s per clip (reported). The SR model never sees the prompt or the audio. It adds texture to what the base model drew, and the paper's own examples say so: hair strands, racket strings, fabric.

What it takes to run

Peak memory for Pro at SD is 72.8 GiB with everything resident (reported). The release streams transformer blocks from CPU memory, two on the GPU at a time. That brings Pro SD to 21.7 GiB, and to 7.9 GiB on the 16 GB preset, which also quantizes the text encoder to NF4. On 16 GB, Pro needs about 58 GiB of host RAM (reported). The paper notes that NF4 is the one setting that changes the output.

Times for one 5-second clip with the full, non-distilled model (reported, Table 10):

RTX 4090RTX 5090H100
Lite SD437 s309 s239 s
Pro SD936 s754 s356 s
Pro Full HD1247 s854 s402 s

That is a quarter of an hour per Pro clip on a 4090. The distilled checkpoint cuts the forward count tenfold against an unaccelerated run. The paper publishes no table for it, so I won't guess its wall time.

How it compares

A radar chart of 14 normalized VABench metrics for LTX 2.5 (two-stage Full HD), Kandinsky 6.0 Video Pro and Kandinsky 6.0 Video Lite. The three polygons overlap closely; Pro sits outside on lip-sync, desync, audio-video alignment and visual realism, LTX 2.5 outside on alignment and expressiveness, and Lite outside on visual QA and audio QA.
VABench, normalized, with desync and WER inverted. The only external model on the chart is LTX 2.5 (paper, Figure 14).

On VABench, Pro leads LTX 2.5 on lip-sync (2.072 vs 1.567) and desync (0.396 vs 0.521, lower is better) (reported). LTX 2.5 keeps the VLM-judged alignment and expressiveness scores and the WER. Two things limit the comparison. LTX 2.5 is the only external model on the chart. And it ran at Full HD while both Kandinsky models ran at 768 x 512. Lite beats Pro on 6 of the 13 VABench rows, and on WER, with a tenth of the parameters (reported, Table 11).

The human side-by-sides use roughly 200 or more comparisons per criterion, majority-voted (reported). Against LTX 2.5, Pro wins overall visual quality 0.49 to 0.26 and speech quality 0.33 to 0.16. Audio-video sync is 82% ties (reported, read off the chart).

Stacked bars for ten criteria, Kandinsky 6.0 Video Pro vs LTX 2.5 in text-to-audio-video. Overall visual quality 0.49 vs 0.26; artifacts 0.34 vs 0.25; motion realism 0.21 vs 0.13; visual prompt following 0.31 vs 0.23; camera motion 0.20 vs 0.15; speech quality 0.33 vs 0.16; audio-video sync 0.05 vs 0.13 with 0.82 ties; overall audio quality 0.12 vs 0.21; sound prompt following 0.24 vs 0.28; user task solving 0.44 vs 0.28.
Pro vs LTX 2.5, text-to-audio-video: green is Kandinsky, grey ties, dark blue LTX 2.5 (paper, Figure 17a).

Against closed models the paper is candid. Veo 3.1 Fast wins user task solving 0.48 to 0.23 and sound prompt following 0.34 to 0.13. Kandinsky wins artifacts 0.50 to 0.19 (reported). MiniMax H3 and Seedance 2.0 are "preferred on the majority of criteria" and on visual criteria respectively. Kling 2.6 is "competitive".

Stacked bars for ten criteria, Kandinsky 6.0 Video Pro vs Veo 3.1 Fast in text-to-audio-video. Overall visual quality 0.32 vs 0.42; artifacts 0.50 vs 0.19; motion realism 0.20 vs 0.14; visual prompt following 0.11 vs 0.40; camera motion 0.33 vs 0.12; speech quality 0.10 vs 0.26; audio-video sync 0.05 vs 0.16; overall audio quality 0.05 vs 0.32; sound prompt following 0.13 vs 0.34; user task solving 0.23 vs 0.48.
Pro vs Veo 3.1 Fast, text-to-audio-video: Kandinsky has fewer artifacts and better camera motion; Veo follows the prompt better and sounds better (paper, Figure 18a).

What I could not check

The design is easy to state. Keep the video model you trust, grow a narrow audio model beside it, and let the two attend to each other every block on a shared, approximate clock. What makes the clips talk is training pressure: filtered data, a face loss, and a reward stack that scores intelligibility with the same recogniser the paper then reports. Read the 47% with that in mind, and the 28% next to it.

Update: FastVideo support

Added 7 October 2026.

A day after release, Hao AI Lab's FastVideo announced Kandinsky 6 support: Pro 30B and Lite 3B, each with "a 10-step distilled pi-Flow version, plus video super-resolution up to 4x," and a recipe page. I read the port in a shallow clone of FastVideo at commit 33d8173 and the headers of the Diffusers checkpoints it loads. It is a careful port with honest documentation, and it is also the clearest place I've found to see what π-Flow does at inference time, so most of this section is about that.

FastVideo's Kandinsky 6 recipe picker. Six recipe cards: Kandinsky 6 Pro TI2VA, Pro pi-Flow, Lite TI2VA, Lite pi-Flow, VSR and VSR distilled. The runtime row offers NVIDIA CUDA with 1 GPU configured and the GPU model not recorded. Below, the Pro TI2VA card lists the model kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers and the command python examples/inference/basic/basic_kandinsky6_ti2va.py.
The six recipes. Note the runtime line: one GPU configured, GPU model not recorded (FastVideo Kandinsky 6 cookbook, screenshot 7 October 2026).

What was ported

Six recipes, all on one GPU: Pro and Lite, each as the 50-step base (guidance 5.0) and the 10-step distilled checkpoint (guidance 1.0), plus the two super-resolution models. FastVideo loads Kandinsky Lab's own kandinskylab/Kandinsky-6.0-*-Diffusers repositories directly, with no conversion and no re-hosted weights, and its local parity test compares its transformer against the Diffusers reference class (Kandinsky6Transformer3DModel), not against the original kandinsky-6 repository this article read.

I range-read those Diffusers checkpoints to make sure they are the same models. The distilled Pro transformer is 30,139,423,760 parameters, the same count as above. The distilled Lite, which I had not counted before, is 3,177,988,112. That is 1,355,688 more than base Lite, and the difference is exactly the widened output heads: the video head goes from 64 to 640 outputs over a 1,792-wide model (1,032,768 extra parameters with bias) and the audio head from 40 to 400 over 896 (322,920). Pro's difference checks the same way, to the parameter. So the distilled models are the base architecture with ten times the output, nothing more. FastVideo's config describes the same heads as out_visual_dim 160 and out_audio_dim 400: 16 latent channels times 10 before the 2x2 spatial patch, and 40 audio channels times 10.

FastVideo is upfront about what it left out, and the list is worth reading before you choose it over the original code. MagCache, which skips 51 of Pro's 100 forwards in the original, is parsed and dropped. NABLA sparse attention is wired but unverified, so attention is dense (Kandinsky's Diffusers pipeline doesn't enable NABLA either). There's no video-only mode, no prompt-expansion pass, and image conditioning supports only the tail_cond_first_frame scheme. Seeds don't match Diffusers: same number, different noise. The default size is 768 x 512, the resolution of the VABench runs above, not the base model's 864 x 480 maximum.

How π-Flow spends ten calls

A normal flow-matching sampler asks the network for one velocity per step and takes one Euler step with it. To get from noise to a clip in 10 steps that way, each step has to jump a tenth of the trajectory on a single straight-line guess, which is where few-step samplers go blurry. π-Flow (Chen et al.) changes what the network returns. FastVideo's scheduling_piflow.py shows how it is consumed.

Each call returns n_grid = 10 predictions of the clean sample, x0, per channel, spread evenly across the segment of time the step has to cover. The scheduler then builds a "network-free DX policy": at any time inside the segment it linearly interpolates between the two nearest x0 predictions and turns that into a velocity, (x_t - x0) / σ. It integrates that velocity with small Euler substeps, 128 per unit of raw time, without touching the network. So one forward pass buys a curved path through its segment instead of one straight step. That is why the output heads are ten times wider. The distilled model also runs without classifier-free guidance: FastVideo, like the Diffusers pipeline, raises an error unless guidance is exactly 1.0, so each of the 10 calls is one forward pass, not two.

raw time 1 (noise)0 (clean)σ 1.00σ 0.98σ 0.95σ 0.92σ 0.87σ 0.82σ 0.74σ 0.64σ 0.48σ 0.22network call (10 x0 predictions each)policy substep (interpolate, then one Euler step; no network)
10 network calls · guidance 1.0 · 10 forwards
124 policy substeps in total
last segment is 0.5× the others (7 substeps vs 13)
Computed with the formulas in FastVideo's PiflowScheduler. The distilled checkpoints were trained for 10 calls (VSR: 2); the scheduler accepts any count, but other counts are off the training distribution. σ is the shifted noise level each call sees.

Working the scheduler's formulas for the shipped settings: with 10 calls and final_step_size_scale 0.5, each segment covers 1/9.5 of raw time and the last covers half that. The policy takes 13 substeps per full segment and 7 in the last, 124 in all. Those substeps are an interpolation and an axpy over the latent, which is nothing next to a 30B forward pass. The super-resolution model uses the same scheduler with a shift of 3.5 and 2 calls, for 85 and 43 substeps. The scheduler will accept any number of calls, which the docs point out, but the checkpoints were trained for 10 (and 2); other counts are a schedule the student never saw.

Super-resolution, and the part that's bigger than the model

The SR pipeline takes any clip, not just Kandinsky's, at 2x, 2.25x (a 1.125x pixel pre-scale, then 2x, the Full HD path described above) or 4x, processes up to 121 frames at 24 fps, and keeps the source audio. Base SR runs 4 DiT calls per tile, distilled SR 2.

The headers held one surprise. The SR transformer is the 1.41B model from earlier, but the convolutional latent upscaler bank that sits in front of it is 3,654,021,952 parameters: a 2.21B network for 4x and a 1.45B one for 2x, both in BF16. The SR VAE adds 1.74B more in F32. In the distilled SR repository the upscaler file is 7.3 GB against 3.2 GB for the transformer. The "1.41B SR model" is the smallest of the three networks that do super-resolution, which matters if you're sizing memory for it. And the SR configs request NABLA attention for 512-pixel tiles, which FastVideo doesn't wire, so it runs dense and logs a warning.

What it costs to run

FastVideo's docs give the Pro DiT as 30.1B parameters, 60 GB in BF16, and the shared Qwen2.5-VL text encoder as 16.6 GB; my header count of 8,292,166,656 BF16 parameters agrees. Lite's DiT is 6.4 GB. FastVideo has no NF4 text encoder or two-blocks-on-GPU streaming like the original's 16 GB preset; its examples keep the DiT on the GPU and offload the text encoder, so for Pro that means a card with room for 60 GB of weights, or its generic DiT offload.

I couldn't find a single timing. FastVideo's own evidence note says every recipe is "Source-backed" by static validation only: "No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are Unknown and deliberately not claimed." I respect that more than a number with no setup attached. The original paper's table above remains the only wall-clock data for these models, and it covers only the non-distilled ones. For FastVideo's work on another audio-video model, including what its quantized builds look like on consumer cards, see FastH3.

Update sources: FastVideo at commit 33d8173, specifically fastvideo/models/schedulers/scheduling_piflow.py, fastvideo/models/dits/kandinsky6.py, docs/inference/kandinsky6.md, docs/inference/kandinsky6_sr.md, docs/cookbook/kandinsky6.md, docs/assets/cookbook-recipes.json and tests/local_tests/kandinsky6/README.md; the config.json, scheduler configs and safetensors headers of kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers, -Lite-distill-5s-Diffusers, -Lite-5s-Diffusers, -VSR-5s-Diffusers and -VSR-distilled2steps-5s-Diffusers, read by HTTP range request; and the π-Flow paper.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Kandinsky 6.0 Video: two token streams on one clock, and where the lip-sync comes from", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026kandinsky6video,
  author = {Satyajit Ghana},
  title  = {Kandinsky 6.0 Video: two token streams on one clock, and where the lip-sync comes from},
  url    = {https://ai.thesatyajit.com/articles/kandinsky-6-video},
  year   = {2026}
}
share