# CSFM: flow matching never had to start from Gaussian noise

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/csfm-flow-matching
> date: 2026-10-02
> tags: flow-matching, diffusion, image-generation, generative-models, representation-learning, explainer

Flow matching has one degree of freedom almost nobody spends: the source distribution. You pick the velocity field, you pick the target data, and then — out of habit inherited from diffusion — you start every sample from the same fixed Gaussian, $\mathcal{N}(0, I)$, the same cloud of noise for a photo of a dog and a photo of a giraffe. The flow-matching framework does not require this. It places no restriction on the source at all.

CSFM — Condition-dependent Source Flow Matching, from NYU and KAIST, at NeurIPS 2026 ([*Better Source, Better Flow*](https://arxiv.org/abs/2602.05951), Kim et al.) — takes that unused freedom and learns the source. Instead of one Gaussian for everything, it trains a small network that reads the text condition and produces a condition-dependent source distribution $p_\phi(X_0 \mid C)$, jointly with the velocity field, under the same flow-matching loss. The headline is convergence: the paper reports up to **3.01x faster FID convergence** and **2.48x faster CLIP-Score convergence** on captioned ImageNet-1K, in a structured latent space. Nothing about the backbone changes. The win comes from where the samples start.

This is a mechanism I can explain from first principles and then check against the paper's own tables, which is the kind of claim worth reading the arXiv HTML for. Let me build up to why the source matters, then audit the three numbers that carry the argument.

## What flow matching is, and why the source got fixed

Flow matching learns a continuous-time velocity field $v_\theta(x, t)$ that transports a source distribution $p_0$ into a target distribution $p_1$ by integrating an ODE:

$$\frac{d}{dt}X_t = v_\theta(X_t, t), \quad t \in [0, 1].$$

Training is almost embarrassingly direct. Take a source sample $X_0$ and a data sample $X_1$, draw them as a pair (a *coupling* $\pi$), lay a straight line between them,

$$X_t = (1 - t)\,X_0 + t\,X_1, \qquad \Delta := X_1 - X_0,$$

and regress the network onto the displacement $\Delta$ at a random time $t$:

$$\mathcal{L}_{\mathrm{FM}}(\theta) = \mathbb{E}_{t,\,\pi(X_0, X_1, C)}\big[\lVert v_\theta(X_t, t, C) - \Delta \rVert^2\big].$$

Each training pair follows a straight line, but the field the network ends up learning is the *average* displacement over every pair passing through a given point, $u_t(x) = \mathbb{E}[\Delta \mid X_t = x]$, and that average is generally curved — because many pairs, with different $\Delta$, cross the same point at the same time and pull the average in different directions.

The source $p_0$ is usually a standard Gaussian for two reasons, neither of them principled. It is what diffusion models use, and under the standard coupling it is convenient: draw noise, draw data, pair them independently. But a Gaussian carries no information about the target. For conditional generation that is a waste, because the condition $C$ is right there and could shape the source too, not just modulate the velocity network. This is the same unused lever that [the Diffusion Transformer](/architectures/diffusion-transformer) write-up traces through the move from noise-prediction to rectified flow: the interpolant is a modelling choice, and so are its endpoints.

## The one term that the source controls

Why would moving the source help at all? The answer is a two-line decomposition of the loss. The flow-matching objective splits exactly into two pieces (the paper derives it in Appendix C):

$$\mathcal{L}_{\mathrm{FM}} = \underbrace{\mathbb{E}\big[\lVert v_\theta(X_t, t) - u_t(X_t) \rVert^2\big]}_{\text{approximation error}} \;+\; \underbrace{\mathbb{E}\big[\mathrm{Var}(\Delta \mid X_t)\big]}_{\text{intrinsic variance}}.$$

The first term is the model's fault — how well the network fits the true marginal field. The second term is not the model's fault at all. It depends only on the coupling $\pi$, not on $\theta$: it is the variance of the displacement $\Delta$ among all the pairs that pass through a given point $(x, t)$. Train forever, grow the network without bound, and the intrinsic variance does not move, because no network can predict a single velocity at a point where the true targets genuinely disagree.

Here is the geometric reading, and it is the whole paper in one sentence: for straight-line interpolants, the intrinsic variance vanishes exactly when the interpolants do not intersect. Two paths crossing at a point means two different $\Delta$ values supervising the same $(x, t)$ — conflicting instructions, averaged into a curved, hard-to-fit field. Fewer crossings means cleaner supervision, a straighter field, and faster convergence.

And the coupling — hence the crossings, hence the intrinsic variance — is exactly what the source distribution controls. A fixed Gaussian far from the data, shared across every condition, produces paths that fan out across the whole space and cross constantly. A source that already sits near the condition's target produces short, local, mostly parallel paths that barely cross. That is the lever.

<Figure
  src="https://ai.thesatyajit.com/articles/csfm-flow-matching/fig1.png"
  alt="Two panels. Left, Standard FM: a fixed source plane on the left, arrows crossing each other as they transport to clustered targets on the right. Right, CSFM: a learned source plane whose clusters are already sorted by caption, with arrows running almost straight across to the matching target clusters."
  caption="Standard flow matching transports a single fixed source to every target, so the paths cross; CSFM's learned source is sorted by condition, so the paths run nearly straight to their matching targets. (CSFM, Figure 1.)"
/>

## Watch the source change the crossings

The decomposition is abstract, so here is the coupling you can poke at. Six conditions, each with a target cluster on a ring. Switch the source between the fixed Gaussian everyone inherits and a learned, condition-shaped source, and watch two exact quantities: how many interpolants cross (the intrinsic-variance proxy — it vanishes precisely when the count hits zero), and how many Euler steps it takes to integrate a sample to the data.

<SourceFlow />

The velocity field in the widget is exact and closed-form for these Gaussians, so every sampler error you see is the integrator's, not a model's. With the fixed source, the interpolants fan out of one blob at the origin and cross hundreds of times; a sample dropped into the ambiguous middle has to be pushed through a bent field and needs ten-plus Euler steps to land. With the learned source relocated next to each target, the paths are roughly half as long, cross several times less often, and a handful of Euler steps — sometimes one — lands on the data. That is the mechanism CSFM is after. The measured image numbers are below; the toy only shows the shape of the argument.

## Learning the source without it collapsing

The idea is one substitution: replace the condition-agnostic Gaussian in the loss with a learnable, condition-dependent source, and expose the induced coupling to the objective rather than fixing it.

$$\pi_\phi(X_0, X_1, C) = p_\phi(X_0 \mid C)\, p_1(X_1, C), \qquad X_0 \sim p_\phi(X_0 \mid C).$$

A source generator $g_\phi$ maps the condition to the parameters of $p_\phi$, and $\phi$ trains jointly with $\theta$ under the same flow-matching loss. Simple — and it collapses in two different ways if you are not careful. CSFM's actual contribution is the two things that stop the collapse.

**Collapse one: the source shrinks to a point.** The fastest way to reduce intrinsic variance is to drive the source variance $\sigma_\phi^2(C)$ to zero — a deterministic source has no internal spread to create crossings. But an overly concentrated source has almost no support, and a flow model cannot recover a full target distribution from a near-point source; in the paper's toy, it fails outright. The source must stay spread out. CSFM parameterizes it as a conditional Gaussian

$$p_\phi(X_0 \mid C) = \mathcal{N}\big(\mu_\phi(C),\, \mathrm{diag}(\sigma_\phi^2(C))\big)$$

and then regularizes **only the variance**, pulling $\sigma_\phi^2(C)$ toward one while leaving the mean $\mu_\phi(C)$ completely free:

$$\mathcal{L}_{\mathrm{VarReg}}(\phi) = \mathbb{E}_C\Big[D_{\mathrm{KL}}\big(\mathcal{N}(\mu_\phi(C), \mathrm{diag}(\sigma_\phi^2(C))) \,\Vert\, \mathcal{N}(\mu_\phi(C), I)\big)\Big].$$

The design choice is the thing inside that KL: the reference distribution has mean $\mu_\phi(C)$, not $0$. Standard VAE-style regularization pulls the source toward $\mathcal{N}(0, I)$, which anchors the *mean* at the origin too — and that is exactly the mobility the source needs. If the mean cannot move, the source cannot relocate next to the target, and you are back to entangled paths. Variance-only regularization keeps the support alive while letting the mean go where it wants.

<Figure
  src="https://ai.thesatyajit.com/articles/csfm-flow-matching/fig2.png"
  alt="A grid of transport-trajectory plots for two toy datasets (Eight Gaussians, Two Moons) under five source designs: fixed standard Gaussian with entangled paths; deterministic mapping that fails to cover the target; conditional Gaussian that collapses; conditional Gaussian with standard KL whose paths are still entangled; and conditional Gaussian with variance-only regularization, whose paths are disentangled and target-aligned."
  caption="Source designs on two toy datasets. Fixed Gaussian (A) entangles; a deterministic source (B) and a collapsed conditional Gaussian (C) lose the target; standard KL (D) prevents collapse but pins the mean and stays entangled; variance-only regularization (E) frees the mean and disentangles the paths. (CSFM, Figure 2.)"
/>

**Collapse two: the source gets no gradient.** Modern text-to-image backbones — MM-DiT, the Lumina-style UnifiedNextDiT, the kind of unified-sequence conditioning that [Qwen-Image-2.1](/articles/qwen-image-2-1) uses — integrate the condition so tightly into the velocity network that the network can explain most of the conditional structure by itself. That leaves almost no learning signal flowing back to the source generator through the flow-matching loss, so in practice the source barely trains. CSFM adds an explicit, source-specific supervision: a directional alignment loss that asks the source sample to point the same way as its target, without matching magnitudes or shrinking the spread,

$$\mathcal{L}_{\mathrm{align}}(\phi) = \mathbb{E}\Big[\,1 - \frac{X_0 \cdot X_1}{\lVert X_0 \rVert\,\lVert X_1 \rVert}\Big],$$

and the final objective is just the three terms added up, $\mathcal{L}_{\mathrm{total}} = \mathcal{L}_{\mathrm{FM}} + \lambda_{\mathrm{VarReg}}\mathcal{L}_{\mathrm{VarReg}} + \lambda_{\mathrm{align}}\mathcal{L}_{\mathrm{align}}$. Cosine, not MSE, is deliberate: matching $\ell_2$ distance over-constrains the source and costs performance, the same failure mode as pinning the mean.

## Checking the ablation

The component table is where the two regularizers earn their place, on captioned ImageNet-1K at 256x256, RAE latents, 100K iterations, 50-step Euler, no guidance. These are the paper's reported numbers; I am reading them off Table 1, not re-running the training.

| Source | Regularizer | FID (lower better) | CLIP | FDD (lower better) |
|---|---|---:|---:|---:|
| Fixed $\mathcal{N}(0, I)$ | — | 3.036 | 0.3398 | 69.08 |
| Learned $p_\phi(X_0 \mid C)$ | none | NaN (collapse) | — | — |
| Learned | standard KL | 2.904 | 0.3405 | 65.92 |
| Learned | variance-only | 2.765 | 0.3404 | 60.64 |
| Learned | MSE align + VarReg | 2.942 | 0.3410 | 69.89 |
| Learned | cosine align + VarReg | **2.453** | 0.3420 | 54.00 |

Everything the method claims is visible here. With no regularizer, the source collapses and the run goes to NaN. Standard KL recovers and even beats the fixed baseline, but only a little (3.036 to 2.904), because it pins the mean. Variance-only regularization does better (2.765) by freeing the mean. Then cosine alignment takes it to 2.453 — and MSE alignment actively hurts (2.942, worse than VarReg alone), confirming that the *direction*, not the distance, is what the source needs.

Two arithmetic observations of my own on those reported numbers. From the fixed baseline to full CSFM, FID drops 19.2% (3.036 to 2.453). And of that total drop, variance regularization delivers a touch under half — (3.036 − 2.765) is about 47% of (3.036 − 2.453) — with cosine alignment supplying the rest. The alignment loss is not a cosmetic add-on; it is roughly half the win. The paper also reports a parameter-matched fixed-Gaussian baseline (FID 2.925) built to absorb the extra parameters the source generator adds, and CSFM still beats it by 16.1%, so the gain is not simply "more parameters."

## Checking the few-step robustness

Straighter flows should survive coarse integration, so the second claim is about few-step sampling. If the field is straight, dropping from 50 Euler steps to a handful barely hurts; if it is bent, it falls apart.

<Figure
  src="https://ai.thesatyajit.com/articles/csfm-flow-matching/fig3.png"
  alt="Two line charts of FID against sampling steps. In both, CSFM (purple) sits below FM (grey) and the gap widens as steps drop toward 2. Left panel is plain Flow Matching; right panel is after 1-Reflow, where CSFM stays nearly flat from 50 down to 2 steps while FM rises sharply."
  caption="FID versus sampling steps for standard FM and CSFM, plain (left) and after one round of reflow (right). CSFM degrades far more gracefully toward 2 steps, the signature of a straighter transport field. (CSFM, Figure 4.)"
/>

Two measurements, both reported. Straight out of the box, cutting 50 steps to 3 costs **8.75 FID** for CSFM against 12.47 for the fixed-source baseline. After one round of reflow — fine-tuning the flow on its own sampled pairs for 20K steps, with the source generator frozen — the gap is much larger: going 50 steps to 2 costs CSFM only **3.51 FID**, while standard FM loses **11.75**. CSFM's degradation is about 70% smaller. The learned source does not just converge faster during training; it leaves behind a field that reflow can straighten into something genuinely few-step. This is the complement of the point in [the few-step Qwen-Image work](/articles/qwen-image-2-1-few-step): curved paths are what force small Euler steps, and the cheapest way to earn back steps is to stop bending the path in the first place.

## When it works: the target has to be structured

The honest part of the paper is a boundary condition. Learning the source is not free money; its benefit depends entirely on the geometry of the *target* latent space, and CSFM shows it both ways.

The default target here is an RAE — a representation autoencoder whose latent is a frozen self-supervised feature space (DINOv2, or SigLIP2 at scale) with a trained decoder, so the latent is semantically organized: images with the same caption land in a compact cluster. Compare that to an SD-VAE latent, optimized for reconstruction, where the same images are smeared across the space. The difference matters because a condition-dependent source works by placing a mean $\mu_\phi(C)$ near the target for $C$ — and that mean is only well-defined if the target for $C$ is compact. In a structured RAE space it is; CSFM's t-SNE shows the learned source inheriting the target's class structure. In an entangled SD-VAE space the target for a caption is spread across distant modes, the source mean is ambiguous, the learned source behaves like a fixed Gaussian, and the gain largely evaporates. The representation-autoencoder idea itself — run the generative flow in a feature space that already carries meaning — is the same bet [TencentARC's GAE](/articles/gae-geometry-native) makes for geometry, and the same tension [Sana](/articles/sana)'s deep-compression autoencoder lives with: the autoencoder decides how much structure the flow has to invent.

<Callout type="warning">
This is the failure mode to remember: CSFM is a bet on the autoencoder. In a structured latent (DINOv2, SigLIP2) the source learning pays off; in an entangled one (SD-VAE) it mostly does not. The paper is explicit that when a single condition maps to a highly multimodal target, the learned source collapses back toward a plain Gaussian prior and buys little.
</Callout>

## Checking it at scale

Toy-to-ImageNet is one thing; the last table asks whether it survives a real text-to-image model. CSFM scales the default recipe to a 1.3B-parameter UnifiedNextDiT, swaps in a Qwen3-0.6B text encoder for longer prompts, pretrains on roughly 36M BLIP3o samples, and evaluates on GenEval and DPG-Bench.

| Model (1.3B) | GenEval | DPG-Bench |
|---|---:|---:|
| Standard FM | 0.77 | 78.31 |
| CSFM | **0.80** | **81.11** |

GenEval goes 0.77 to 0.80, DPG-Bench 78.31 to 81.11 — consistent, same direction as the small-scale result, and the paper is candid that these benchmarks are saturating at this scale, so the quantitative gap understates what the qualitative samples show. It is a real but modest bump, which is the right thing to report rather than a hero number.

## What it costs, and the licence

The costs are stated plainly in the paper and worth keeping in view. CSFM adds a source generator (extra parameters, though the parameter-matched baseline shows the win is not just capacity) and two loss terms with two more hyperparameters to balance. It uses AutoGuidance rather than classifier-free guidance, because a learned source has no natural unconditional branch. It is validated only on text-to-image; other modalities and frontier-proprietary scale are future work. And the gain is representation-dependent in the way above — a method that helps a lot or barely at all depending on an autoencoder choice is a method you have to profile before you adopt.

One thing the paper does not mention and the repository settles: the official PyTorch code at `github.com/junwankimm/CSFM` ships with **no licence file**, which under default copyright means all rights reserved — reading it is fine, building on it is not, until that changes. The caption dataset, `junwann/CSFM-ImageNet1K-Caption` on Hugging Face, is tagged MIT. If you want to use this, the dataset is reusable today and the code is not.

The lasting idea is smaller than the benchmark table and more useful: in conditional flow matching the source is a first-class design choice, not a constant. Fixing it to Gaussian was a convenience borrowed from diffusion, and it was leaving the intrinsic-variance term — the part of the loss no amount of training can touch — larger than it needed to be. CSFM's contribution is less "learn the source" (others tried) than the recipe that makes learning it stable: regularize the variance, not the mean; align the direction, not the distance; and only expect it to pay off when the target space is structured enough for a condition to name a place.
