2026-10-06 · 27 min · video-generation · audio · diffusion · flow-matching · reinforcement-learning · distillation
Why read this
Notabletop 60%Explains joint audio-video denoising from the code (230:1 tokens, RoPE scaled 0.144) and shows the 47% WER cut is graded by its own reward model.
- Original analysis
- Runs on a consumer GPU
- Concrete numbers to act on
Image & video generationMITPractitioner model
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 2 of 3: Open code or weights with real limits
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 59 of 100, ranked 234 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
Kandinsky Lab released Kandinsky 6.0 Video this week: two diffusion models that write a 5-second clip and its soundtrack in one pass. Lite is advertised at 3B parameters, Pro at 29B. Both take text, or text plus a first frame, and return SD video with 44 kHz audio, speech and lip movement included. A separate super-resolution model takes the picture to 1920 x 1080. Weights, code and a diffusers integration are MIT-licensed. The paper is unusually complete. It gives data filters, per-stage learning rates, reward weights and timings on seven GPUs.
This piece explains how a model draws a picture and its sound together, from the token up. Then it
checks the release against what can be read without a GPU. I read the paper, the
kandinskylab/kandinsky-6 code at commit ea44d1b, the
diffusers configs, and the safetensors headers of the three ungated transformers, over HTTP range
requests. The non-distilled Pro checkpoint is gated behind a click-through. Its parameter count comes
from the Hub API, which reads the same header. I ran nothing on a GPU. Every number below is labelled
measured (I read it from a file), reported (the paper's figure) or reasoned (my arithmetic on
the other two).
- task
- image-to-video
- library
- diffusers
- license
- mit
- safetensors
- 10 shards
- largest file
- 69.73 GB
- files
- 40
- downloads
- 173
- likes
- 15
- gated
- auto
Transformer only: 30,136,326,248 parameters, 4,730,755,072 of them stored in F32 (measured, Hub API), which is why the file is 69.7 GB. The distilled Pro is 30,139,423,760 parameters, 60.3 GB, almost all BF16 (measured from its header). Gated click-through; MIT licence.
repo last modified 2026-10-06
- task
- image-to-video
- library
- diffusers
- license
- mit
- safetensors
- 10 shards
- largest file
- 7.48 GB
- files
- 40
- downloads
- 53
- likes
- 9
Transformer only: 3,176,632,424 parameters (measured from the header). The repo also ships the Qwen2.5-VL-7B and CLIP text encoders, the Hunyuan video VAE, and the MMAudio VAE and vocoder. Ungated; MIT licence.
repo last modified 2026-10-06
- license
- MIT
- branch
- main
- tests
- 2 files
- source
- 631.3 kB
- commit date
- 2026-10-06
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-06 at ea44d1b — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history

Two ways to put sound on a picture
The obvious way is a cascade. Generate the video, then hand it to a video-to-audio model that listens to the pixels and writes a soundtrack. MMAudio, whose autoencoder Kandinsky 6.0 reuses, is that second stage. A cascade is modular, and each half can be trained on its own data. It also gets one thing structurally wrong. Speech has to drive the mouth, and a cascade can only drive it the other way. By the time the audio model runs, the lips are frozen pixels. The audio can only chase them.
The joint way denoises both at once. Every step, the video sees where the audio is going and the audio sees where the video is going, so a word and the jaw that shapes it are one decision. There are two designs. One puts everything into a single token sequence with modality-specific weights, which is what MAGI-2 does. The other keeps two transformers, one per modality, and has them talk through cross-attention in every block. Kandinsky 6.0 is the second kind, like Ovi and LTX-2. I took apart LTX-2's version of that block in the restore and relight LoRAs piece. MiniMax's H3, which FastH3 and FreeVideo work on, is another large joint audio-video transformer.
One block, two streams

Read in the code (FusedTransformerDecoderBlock in dit.py), one block does this, in order:
- Self-attention, per stream. Video tokens attend to video tokens with a 3D RoPE over (time, height, width). Audio tokens attend to audio tokens with a 1D RoPE over time.
- Text cross-attention, per stream. Each stream has its own small text transformer over the Qwen2.5-VL embedding, so the video and the audio read the same prompt through different weights.
- Cross-modal attention, both ways. Video queries attend to audio keys and values, and audio
queries attend to video. The keys and values are taken from the other stream's state after
self-attention but before its text cross-attention, and each result is added back through a
learned gate (measured,
dit.py). - Feed-forward, per stream.
Each stream gets its own timestep embedding. During joint training the noise level is drawn independently for video and audio, "so that the two modalities are noised to different levels" (reported). This is what lets one network also do audio-to-video and video-to-audio: hold one stream at zero noise and denoise the other.
The two streams are deliberately unequal. The paper gives the split in round numbers. Here it is from the headers:
| Lite (paper) | Lite (header) | Pro (paper) | Pro distill (header) | |
|---|---|---|---|---|
| video stream | 2B | 1,909,391,360 | 19B | 18,375,260,160 |
| audio stream | 0.6B | 543,657,984 | 5B | 4,909,455,360 |
| cross-attention | 0.4B | 595,152,896 | 5B | 5,664,921,600 |
| text blocks and embeddings | not counted | 128,430,184 | not counted | 1,189,786,640 |
| total | 3B | 3,176,632,424 | 29B | 30,139,423,760 |
The header columns are measured. I bucketed every tensor by its prefix. "Cross-attention" means the
va_ and av_ attention and modulation tensors in every block. The total row of the paper is the sum
of its own three lines; it leaves out the per-stream text transformers, which is most of the 1.1B gap
on Pro (reasoned). The cross-attention is 18.7% of Lite and 18.8% of Pro (reasoned from the measured
counts). Lite has 32 blocks with a 1,792-wide video stream and an 896-wide audio stream. Pro has 60
blocks at 4,096 and 2,048 (measured, transformer/config.json).
Counting one clip
Now the token budget, because it explains most of the design. Video goes through the Hunyuan VAE,
which compresses 8x in space and 4x in time into 16 channels. The transformer then patches 2 x 2. One
video token is a 16 x 16 pixel square of one latent frame. The default clip is 121 frames at 24 fps,
which the causal VAE turns into 31 latent frames: the first frame alone, then four per latent. At
864 x 480 that is 54 x 30 = 1,620 tokens per latent frame, and 50,220 video tokens per clip
(reasoned from the configs and pipeline.py).
Audio goes through MMAudio's 44.1 kHz autoencoder, which turns 1,024 samples into one 40-channel
latent frame. The DiT does not patch it: each latent frame is one token. The pipeline sizes the audio
to the video as ceil(121 / 24 x 44,100 / 1,024) = 218 tokens (measured, prepare_latents.py). That
is one audio token per 23.2 ms against one video latent frame per 166.7 ms, about 7.2 audio tokens
under each latent frame (reasoned).
So a clip is 50,220 video tokens next to 218 audio tokens, about 230 to 1. The audio stream holds 16% of Pro's weights and sees 0.4% of the tokens. Counting each stream's weights against its own tokens, audio is about 0.12% of the linear-layer work per step (reasoned). The audio stream is cheap to make wide. The cross-attention is where the two sizes meet: 50,220 x 218 query-key pairs each way, against 2.5 billion pairs in video self-attention.
SD landscape, the release default
video stream 18.38 G · audio stream 4.91 G · cross-attention 5.66 G parameters
Pick a video latent frame to see which audio tokens play under it, and where RoPE puts them.
How does a video token know which audio tokens are simultaneous with it? Through RoPE. The model
passes rotary positions into the cross-modal attention too (ca_rope: true). Video queries carry their
3D position; audio keys carry their 1D one. The paper says the video frame index is normalized to 24
fps, and the audio RoPE's frequencies are all multiplied by audio_freqs_scaling = 0.144 (measured,
transformer/config.json and rope.py). That squeezes 218 audio positions into roughly the range of
31 video positions: the last audio token lands at 217 x 0.144 = 31.2 (reasoned).
Two details in the code make this a hint, not a lock (reasoned, from rope.py):
- The scale is close, not exact. A video latent frame is 4/24 s and an audio token is 1,024 / 44,100 s, so a time-exact scale would be 0.1393. At 0.144, audio tokens at the end of the clip carry positions about one latent frame ahead of the video tokens they play under. The widget shows both.
- The two RoPEs do not share a frequency table. In Lite's 64-channel heads, the video's time axis owns 16 channels at 8 frequencies, while the audio's 1D RoPE spreads 32 frequencies over all 64 channels. Only the first channel pair rotates in matching units. The rest of the audio head lines up with the video's height and width channels.
Neither is a bug. A cross-attention can learn to read a consistent, slightly off signal, and the sync results below suggest this one does. But the alignment is learned, not built in.
The first-frame mode reuses the same machinery. In image-to-audio-video the reference frame's latent is
appended along time, flagged by a mask channel and a learned "token role" embedding. The paper's
ablation found the RoPE position does the binding. Give the reference a later temporal position and the
model uses it as that later frame. The mask separates conditioning from noise. The role embedding is
"partially redundant" (reported). The release code gives the appended frame position 0 (measured,
pipeline.py).
Where 44 kHz comes from
"44 kHz audio" is the sample rate of MMAudio's autoencoder: 44,100 Hz. That stack is a mel-spectrogram
VAE plus a BigVGAN vocoder. The vocoder's upsampling rates multiply to 8 x 4 x 2 x 2 x 2 x 2 = 512
samples per mel frame, over 128 mel bins (measured, vocoder/config.json). The VAE halves the frame
rate again to reach 1,024 samples per latent (reasoned from the two configs). The DiT never sees a
waveform. It denoises 218 x 40 numbers; the audio VAE decodes them to a mel spectrogram and the vocoder
turns that into sound. Audio fidelity is therefore bounded by a frozen, borrowed codec. The paper's
contribution is the transformer that writes into its latent space.
Training: two towers, then a merger

The video stream starts from a Kandinsky 5.0 checkpoint. The audio stream starts from random weights and learns sound alone first, on 40 million audio tracks of 5 to 20 seconds (reported). Then the newly initialised cross-attention layers connect them. Lite trains jointly for 57.5k steps and Pro for 49k, on 7 million audio-video segments. The last 4k to 5k steps put 25% of samples in image-conditioned mode (reported). A batch can also hold audio only, video only or images only (reported).
There is no lip-sync module
Nothing in the architecture is about faces. Lip-sync is pressed in from four directions, all reported:
- Data. A joint clip with a person in it is kept only if its lip-sync alignment score is at least 3 out of 5. A clip with audio-visual events needs a desync score below 0.1. About 25% of the joint data has no faces at all.
- A face-weighted loss. Supervised fine-tuning trains 11 domain experts. Stage 2 is about sync, and its loss is 0.85 x video + 0.15 x audio + 1.0 x a "face loss": the video loss restricted to the facial region in latent space. The 11 experts are averaged into one model, a model soup.
- A lip-sync reward in the RL stage, below.
- Attention-weighted RL. The RL loss weights video patches by how strongly audio tokens attend to them in A2V cross-attention. Faces, hands and instruments get up to 1.5x. The model's own attention decides where sync matters.

The RL stage, and what the 47% measures

The RL recipe is OmniNFT, originally built for LTX-2 and adapted here (reported). For each prompt Pro generates a group of K = 10 clips that differ only in seed (8 for Lite). The sampler is deterministic, DPM2 at 20 steps, so the seed is the only source of variety. Each of eight rewards is centred within the group and scaled by the batch's standard deviation, like GRPO. The weighted advantages then split by modality. Video rewards (HPSv3, VideoAlign) go to video tokens and audio rewards (Audiobox Aesthetics, CLAP, Whisper ASR) to audio tokens. The three sync rewards go to both. Each token gets a reward between 0 and 1. The update is not a policy gradient over log-probabilities, which a flow model does not have cheaply. It re-noises the finished clip, asks the current and previous LoRA to denoise it, and pulls the prediction toward the rollout in proportion to that reward:
where is token 's reward, the rollout, and , the current and previous predictions (simplified: the paper also normalizes each term). A high reward pulls the current model toward the rollout. A low reward pulls its mirror image toward it, which pushes the model away. A small KL term () keeps it near the SFT model. Only a LoRA trains: rank 16, 199M parameters on Pro (0.69%) and 47M on Lite (1.57%) (reported). There is also a hand-tuned schedule that shrinks gradients flowing into the audio stream through V2A attention in some blocks, so video rewards do not wreck the audio.
Now the 47%. The paper's sentence in full: Pro's word error rate on 64 held-out speech prompts, three seeds each, drops "from 0.235 for the SFT model to 0.124, a relative reduction of 47%" (reported, p under 0.001). Lite drops from 0.179 to 0.127, 29% (reported). The arithmetic checks.
The catch is in the next line of the PDF: "WER is computed with Whisper-large-v3, which is also one of the reward models." Speech intelligibility is rewarded twice in the stack. Once directly, as Whisper ASR at weight 0.5. Once inside the lip-sync term, which multiplies the SyncNet score by , where is one minus the WER. Grading an RL policy with the model it was trained against measures how well it learned to please that model. The paper knows this and runs two other checks:
- A held-out sentence set. On 100 Harvard sentences, Pro's WER goes from 0.251 to 0.180, p = 0.005 (reported). That is a 28% cut (reasoned), real but well short of 47%.
- A different recogniser. VABench's speech subset is scored with Qwen3-ASR-1.7B. The RL'd models score 0.103 (Lite) and 0.117 (Pro) against 0.099 for LTX 2.5 (reported). Pro has the highest WER of the three. The paper says the two WERs "are not directly comparable", which is true, and is the point.

The human side-by-side of RL against SFT is the better evidence that RL helped, and it did. Visual quality goes 0.58 to 0.17 for RL, and audio-video sync 0.41 to 0.17 (reported, read off the chart). The text says speech quality "improves significantly" on this benchmark, but the published chart has no speech-quality bar. Audio prompt following is 92% ties. So the RL stage clearly made clips look better and sync better. How much it improved speech rests on a metric graded by its own reward.
Ten steps instead of a hundred
The full Pro samples with 50 Euler steps at guidance 5.0 (measured, pro.yaml). With classifier-free
guidance that is 100 transformer forwards. With MagCache, Pro SD evaluates 49 of those 100 (reported). The
distilled checkpoint, which runs the public demo, uses 10 forwards at guidance 1.0 (measured,
pro-distill.yaml). It gets there in two stages. π-Flow trains a student whose output head
predicts ten velocities at once, a small policy that integrates a segment without calling the network
again. The headers show it: the distilled output layers are 640 wide for video and 400 for audio, ten
times the 64 and 40 of the base heads (measured). Then 300 steps of adversarial fine-tuning (Sim-LADD)
on the student's own rollouts restore the texture that trajectory compression smooths away.
In a human side-by-side on image-to-audio-video, the distilled Pro is preferred 51% to 49% over the full model, and no criterion differs significantly (reported). The demo is not a downgraded model, on the publisher's evidence.
From 480p to Full HD

The base model stops at 864 x 480. Full HD is a separate, text-free pipeline in a different latent space: K-VAE, 64 channels, 16x spatial compression. The clip is pre-scaled 1.125x in pixels to 976 x 544, then encoded once and tiled. A convolutional Latent Upscaler doubles each tile in latent space. The tile is partially re-noised, and a 1.41B-parameter SR-DiT, initialised from Kandinsky 5.0 Video Lite, restores detail. The output is 1952 x 1088, resized to 1920 x 1080 (reported). The released SR-DiT is distilled to two network calls per tile; its checkpoint is 1,414,263,936 parameters (measured, Hub API). On an H100, Full HD from 864 x 480 takes 9 tiles and 39.6 s per clip (reported). The SR model never sees the prompt or the audio. It adds texture to what the base model drew, and the paper's own examples say so: hair strands, racket strings, fabric.
What it takes to run
Peak memory for Pro at SD is 72.8 GiB with everything resident (reported). The release streams transformer blocks from CPU memory, two on the GPU at a time. That brings Pro SD to 21.7 GiB, and to 7.9 GiB on the 16 GB preset, which also quantizes the text encoder to NF4. On 16 GB, Pro needs about 58 GiB of host RAM (reported). The paper notes that NF4 is the one setting that changes the output.
Times for one 5-second clip with the full, non-distilled model (reported, Table 10):
| RTX 4090 | RTX 5090 | H100 | |
|---|---|---|---|
| Lite SD | 437 s | 309 s | 239 s |
| Pro SD | 936 s | 754 s | 356 s |
| Pro Full HD | 1247 s | 854 s | 402 s |
That is a quarter of an hour per Pro clip on a 4090. The distilled checkpoint cuts the forward count tenfold against an unaccelerated run. The paper publishes no table for it, so I won't guess its wall time.
How it compares

On VABench, Pro leads LTX 2.5 on lip-sync (2.072 vs 1.567) and desync (0.396 vs 0.521, lower is better) (reported). LTX 2.5 keeps the VLM-judged alignment and expressiveness scores and the WER. Two things limit the comparison. LTX 2.5 is the only external model on the chart. And it ran at Full HD while both Kandinsky models ran at 768 x 512. Lite beats Pro on 6 of the 13 VABench rows, and on WER, with a tenth of the parameters (reported, Table 11).
The human side-by-sides use roughly 200 or more comparisons per criterion, majority-voted (reported). Against LTX 2.5, Pro wins overall visual quality 0.49 to 0.26 and speech quality 0.33 to 0.16. Audio-video sync is 82% ties (reported, read off the chart).

Against closed models the paper is candid. Veo 3.1 Fast wins user task solving 0.48 to 0.23 and sound prompt following 0.34 to 0.13. Kandinsky wins artifacts 0.50 to 0.19 (reported). MiniMax H3 and Seedance 2.0 are "preferred on the majority of criteria" and on visual criteria respectively. Kling 2.6 is "competitive".

What I could not check
- Output quality. I generated nothing. Every quality claim above is the publisher's, from a human evaluation whose annotator pool and prompt set are not released.
- The non-distilled Pro weights. Gated; the parameter count is the Hub API's reading of its header.
- Whether the RoPE mismatch costs anything. The scale and the frequency layout are measured from the code. Whether a time-exact scale would sync better is an experiment the paper does not report.
The design is easy to state. Keep the video model you trust, grow a narrow audio model beside it, and let the two attend to each other every block on a shared, approximate clock. What makes the clips talk is training pressure: filtered data, a face loss, and a reward stack that scores intelligibility with the same recogniser the paper then reports. Read the 47% with that in mind, and the 28% next to it.
Update: FastVideo support
Added 7 October 2026.
A day after release, Hao AI Lab's FastVideo announced Kandinsky 6
support: Pro 30B and Lite 3B, each with "a 10-step distilled
pi-Flow version, plus video super-resolution up to 4x," and a recipe page.
I read the port in a shallow clone of FastVideo at commit 33d8173
and the headers of the Diffusers checkpoints it loads. It is a careful port with honest documentation, and it is
also the clearest place I've found to see what π-Flow does at inference time, so most of this section is about that.

What was ported
Six recipes, all on one GPU: Pro and Lite, each as the 50-step base (guidance 5.0) and the 10-step distilled
checkpoint (guidance 1.0), plus the two super-resolution models. FastVideo loads Kandinsky Lab's own
kandinskylab/Kandinsky-6.0-*-Diffusers repositories directly, with no conversion and no re-hosted weights,
and its local parity test compares its transformer against the Diffusers reference class
(Kandinsky6Transformer3DModel), not against the original kandinsky-6 repository this article read.
I range-read those Diffusers checkpoints to make sure they are the same models. The distilled Pro transformer is
30,139,423,760 parameters, the same count as above. The distilled Lite, which I had not counted before, is
3,177,988,112. That is 1,355,688 more than base Lite, and the difference is exactly the widened output heads: the
video head goes from 64 to 640 outputs over a 1,792-wide model (1,032,768 extra parameters with bias) and the audio
head from 40 to 400 over 896 (322,920). Pro's difference checks the same way, to the parameter. So the distilled
models are the base architecture with ten times the output, nothing more. FastVideo's config describes the same
heads as out_visual_dim 160 and out_audio_dim 400: 16 latent channels times 10 before the 2x2 spatial patch,
and 40 audio channels times 10.
FastVideo is upfront about what it left out, and the list is worth reading before you choose it over the original
code. MagCache, which skips 51 of Pro's 100 forwards in the original, is parsed and dropped. NABLA sparse attention
is wired but unverified, so attention is dense (Kandinsky's Diffusers pipeline doesn't enable NABLA either). There's
no video-only mode, no prompt-expansion pass, and image conditioning supports only the tail_cond_first_frame
scheme. Seeds don't match Diffusers: same number, different noise. The default size is 768 x 512, the resolution of
the VABench runs above, not the base model's 864 x 480 maximum.
How π-Flow spends ten calls
A normal flow-matching sampler asks the network for one velocity per step and takes one Euler step with it. To get
from noise to a clip in 10 steps that way, each step has to jump a tenth of the trajectory on a single straight-line
guess, which is where few-step samplers go blurry. π-Flow (Chen et al.) changes
what the network returns. FastVideo's scheduling_piflow.py shows how it is consumed.
Each call returns n_grid = 10 predictions of the clean sample, x0, per channel, spread evenly across the segment
of time the step has to cover. The scheduler then builds a "network-free DX policy": at any time inside the segment
it linearly interpolates between the two nearest x0 predictions and turns that into a velocity, (x_t - x0) / σ.
It integrates that velocity with small Euler substeps, 128 per unit of raw time, without touching the network.
So one forward pass buys a curved path through its segment instead of one straight step. That is why the output
heads are ten times wider. The distilled model also runs without classifier-free guidance: FastVideo, like the
Diffusers pipeline, raises an error unless guidance is exactly 1.0, so each of the 10 calls is one forward pass,
not two.
Working the scheduler's formulas for the shipped settings: with 10 calls and final_step_size_scale 0.5, each
segment covers 1/9.5 of raw time and the last covers half that. The policy takes 13 substeps per full segment and 7
in the last, 124 in all. Those substeps are an interpolation and an axpy over the latent, which is nothing next to a
30B forward pass. The super-resolution model uses the same scheduler with a shift of 3.5 and 2 calls, for 85 and 43
substeps. The scheduler will accept any number of calls, which the docs point out, but the checkpoints were trained
for 10 (and 2); other counts are a schedule the student never saw.
Super-resolution, and the part that's bigger than the model
The SR pipeline takes any clip, not just Kandinsky's, at 2x, 2.25x (a 1.125x pixel pre-scale, then 2x, the Full HD path described above) or 4x, processes up to 121 frames at 24 fps, and keeps the source audio. Base SR runs 4 DiT calls per tile, distilled SR 2.
The headers held one surprise. The SR transformer is the 1.41B model from earlier, but the convolutional latent upscaler bank that sits in front of it is 3,654,021,952 parameters: a 2.21B network for 4x and a 1.45B one for 2x, both in BF16. The SR VAE adds 1.74B more in F32. In the distilled SR repository the upscaler file is 7.3 GB against 3.2 GB for the transformer. The "1.41B SR model" is the smallest of the three networks that do super-resolution, which matters if you're sizing memory for it. And the SR configs request NABLA attention for 512-pixel tiles, which FastVideo doesn't wire, so it runs dense and logs a warning.
What it costs to run
FastVideo's docs give the Pro DiT as 30.1B parameters, 60 GB in BF16, and the shared Qwen2.5-VL text encoder as 16.6 GB; my header count of 8,292,166,656 BF16 parameters agrees. Lite's DiT is 6.4 GB. FastVideo has no NF4 text encoder or two-blocks-on-GPU streaming like the original's 16 GB preset; its examples keep the DiT on the GPU and offload the text encoder, so for Pro that means a card with room for 60 GB of weights, or its generic DiT offload.
I couldn't find a single timing. FastVideo's own evidence note says every recipe is "Source-backed" by static validation only: "No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are Unknown and deliberately not claimed." I respect that more than a number with no setup attached. The original paper's table above remains the only wall-clock data for these models, and it covers only the non-distilled ones. For FastVideo's work on another audio-video model, including what its quantized builds look like on consumer cards, see FastH3.
Update sources: FastVideo at commit 33d8173, specifically fastvideo/models/schedulers/scheduling_piflow.py,
fastvideo/models/dits/kandinsky6.py, docs/inference/kandinsky6.md, docs/inference/kandinsky6_sr.md,
docs/cookbook/kandinsky6.md, docs/assets/cookbook-recipes.json and tests/local_tests/kandinsky6/README.md;
the config.json, scheduler configs and safetensors headers of kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers,
-Lite-distill-5s-Diffusers, -Lite-5s-Diffusers, -VSR-5s-Diffusers and -VSR-distilled2steps-5s-Diffusers, read
by HTTP range request; and the π-Flow paper.