~/satyajit

Dust: every token is a member of the population

mdjsonmcp

2026-10-06 · 20 min · explainer · training · pretraining · theory · backpropagation · optimization · transformers

"Backprop has been the only credit assignment algorithm capable of training large neural nets." That is how Samip Dahal opens the thread announcing Dust, from Q Labs, written with Bishwas Mandal, Serdar Gülbahar and Akshay Vegesna. The claim that follows is that Dust is "the first zeroth-order method to pretrain transformers to approach and sometimes even exceed backprop with large amounts of computation", and that it is "on the order of 1,000 to 10,000x more compute efficient than EGGROLL".

Zeroth-order means the optimizer only ever sees loss values. That family has a well-earned reputation for collapsing in high dimensions, so this is worth taking apart.

I read the post, its appendix, and the released code (qlabs-eng/dust, MIT, commit 2fdb01b). I did not run it: the paper configuration wants eight GPUs per run. Numbers below are labelled reported (Q Labs' figure, not re-run), measured (I read or computed it from a file, the code, or the toy on this page), or reasoned (my arithmetic on the other two).

Dust on one layer: jitter every token's output, reward each token's jitter by its own loss change, average the reward-weighted noise into an output error per token, and take the outer product with the input to get the weight gradient. Q Labs' own animation of the method (Dust post, method animation).

What a zeroth-order gradient is

Start with any loss L(w)L(w) over DD parameters. Nudge the parameters in a random direction ε∼N(0,ID)\varepsilon \sim \mathcal{N}(0, I_D) by a small step σ\sigma. To first order,

L(w+σε)−L(w)≈σ⟨∇L,ε⟩.L(w + \sigma\varepsilon) - L(w) \approx \sigma \langle \nabla L, \varepsilon \rangle .

One scalar, the loss change, tells you how much of the gradient lies along ε\varepsilon. Multiply it back onto ε\varepsilon and average over KK directions:

g^=1Kσ∑i=1K(L(w+σεi)−L(w)) εi.\hat g = \frac{1}{K\sigma}\sum_{i=1}^{K} \big(L(w + \sigma\varepsilon_i) - L(w)\big)\,\varepsilon_i .

Because E[εε⊤]=I\mathbb{E}[\varepsilon\varepsilon^\top] = I, the expectation of g^\hat g is the gradient. That is the evolution-strategies estimator of Salimans et al. (arXiv 1703.03864) and the random gradient-free method Nesterov and Spokoiny analysed. Evaluating L(w−σε)L(w - \sigma\varepsilon) as well and using the difference of the pair (antithetic sampling) cancels the even-order terms, so the curvature does not leak into the estimate.

The catch is variance. Each sample is the true gradient's component along one random direction plus everything else that direction carries. For an isotropic Gaussian, the error of g^\hat g has squared norm of roughly (D+1)∥∇L∥2/K(D+1)\lVert\nabla L\rVert^2/K, so the cosine between estimate and truth is about

cos⁡(g^,∇L)≈KK+D.\cos(\hat g, \nabla L) \approx \sqrt{\frac{K}{K + D}} .

To get a cosine of one half you need K≈D/3K \approx D/3 evaluations. Dust's base model has 37.7M parameters (reported). Weight-space ES on it would need about 12.6M forward passes per update for a cosine of 0.5, and at 16,384 forward passes the cosine is about 0.02 (reasoned, from the formula above). That is the wall every weight-space method hits.

EGGROLL (arXiv 2511.16652) attacks the cost per member: it makes each perturbation a rank-rr matrix, so a whole population of perturbed models runs as one batched inference at up to 91% of plain inference throughput (reported, in its abstract). It does not change DD. Every member is still one sequence of the batch, and the population is still bounded by the forward passes you can afford.

Perturb the activations, not the weights

Dust moves the noise. Take one linear layer, yt=Wxty_t = W x_t, at token tt. Backprop's weight gradient for that layer is

∂L∂W=∑tδt xt⊤,δt=∂L∂yt.\frac{\partial L}{\partial W} = \sum_t \delta_t\, x_t^\top, \qquad \delta_t = \frac{\partial L}{\partial y_t}.

The input xtx_t is free: the forward pass already computed it. The only hard part is the output error δt\delta_t, which backprop gets from the chain rule. Dust estimates it by jittering the layer's output instead: yt→yt+σaty_t \to y_t + \sigma a_t with at∼N(0,I)a_t \sim \mathcal{N}(0, I), at every token, in one forward pass. The post writes the reward at token tt as the loss reduction there, plus the decayed reductions at later tokens that the jitter reaches through attention:

rt=∑s≥tγ s−tcs,g^t=−1Kσ∑i=1Krt(i)at(i),G^W=∑tg^t xt⊤.r_t = \sum_{s \ge t} \gamma^{\,s-t} c_s, \qquad \hat g_t = -\frac{1}{K\sigma}\sum_{i=1}^{K} r_t^{(i)} a_t^{(i)}, \qquad \widehat G_W = \sum_t \hat g_t\, x_t^\top .

Here csc_s is the centred loss reduction at token ss, KK is the number of draws (one draw is one independent jitter of all tokens), and γ\gamma is a credit decay that is zero for everything except the attention internals. The last equation is the same outer product backprop forms. Only the source of the error differs. This is node perturbation, which Widrow and Lehr described in 1990 and Werfel, Xie and Seung analysed in 2003; the post cites both.

The textbook argument for node perturbation is dimension: a layer's output has doutd_{\text{out}} numbers and its weights have dout×dind_{\text{out}} \times d_{\text{in}}. The post is candid that this argument, read naively, fails for a transformer. Noise on one sequence is a T×doutT \times d_{\text{out}} tensor, which has at least as many entries as the weight matrix once T≥dinT \ge d_{\text{in}}. If you reward that whole tensor with one scalar, you are back to searching a space as big as the weights.

What rescues it is the per-token reward. Each token's loss mostly reflects its own jitter, so each token's noise is scored by its own number. The search per reward is then doutd_{\text{out}}-dimensional, and the TT tokens of a sequence are TT separate population members evaluated in the same pass. Q Labs calls that a virtual population. A sequence of 2,048 tokens is 2,048 members, where EGGROLL gets one.

Try it: three estimators, same forward passes

The widget below is a toy, not the paper's model. One d×dd \times d linear layer, 32 tokens, a quadratic loss at each token. Three estimators get the same number of forward passes KK:

The two activation estimators rebuild the weight gradient exactly as Dust does, as ∑tg^txt⊤\sum_t \hat g_t x_t^\top. The curve is the cosine to the true gradient.

one linear layer · 32 tokens · same forward passes for all threeruns in your browser
width d
0.000.250.500.751.002481632641282565121k2kforward passes K (log scale)cosine to the true gradient
K = 256
estimatorsearched per rewardcosine at K = 256
weight-space ES (antithetic)256 (d × d weights)0.536
activation noise, one reward per pass512 (T × d activations)0.511
activation noise, one reward per token (Dust)16 (d per token)0.968
A toy, not the paper's model: one d × d linear layer with a quadratic loss at each of 32 tokens, noise scale σ = 0.05, one seeded run per width. Weight-space ES gets antithetic pairs (K/2 of them), which is the kindest version of it. Token losses here do not interact, so the per-token estimator has none of the cross-token interference a real transformer adds through attention.

The toy reproduces the post's dimension argument (measured, one seeded run per width). At width 16, with 256 forward passes, weight ES reaches a cosine of 0.536 and the one-reward activation estimator 0.511, because the second searches 512 numbers per reward against the first's 256. They are equally hopeless. The per-token estimator is at 0.968 with the same passes. At width 64 the gap is wider: at 2,048 passes weight ES sits at 0.442, one reward per pass at 0.557, one reward per token at 0.984. The per-token curve barely moves as the width grows, because what it searches per reward is dd, not d2d^2 and not T⋅dT \cdot d.

The toy is kind in two ways. Its token losses do not interact, so the per-token estimator suffers none of the cross-token interference attention creates in a real model; that interference is what most of Dust's engineering is for. And its loss is exactly quadratic, which flatters antithetic weight ES, since the pair difference is then exactly the directional derivative.

What the released code actually does

The repository is small: dust.py is 607 lines and model.py 201 (measured). The model is the modded-nanoGPT kind of decoder: eight pre-norm blocks of width 512, squared-ReLU MLP, rotary embeddings with normalised queries and keys, value embeddings on alternate blocks, a soft-capped untied head (reported in the appendix, and it matches model.py). It is the nanoGPT speedrun lineage. Nothing in dust.py calls backward(): the model is built with requires_grad_(False) and the update runs under torch.no_grad() (measured).

One update in DUST.step, as I read it:

  1. A clean forward caches every Linear and Embedding's input and output through forward hooks, plus the per-token losses.
  2. Direct draws. direct_errors sets a jitter sigma * a on the chosen sites, reruns the forward, and computes reward = clean_loss - perturbed_loss per token. When several draws share one forward (the batch repeated nrep times), the reward is centred across them; with nrep = 1 it is not. The error estimate accumulates -reward * a / sigma. These rewards use the token's own loss only, so γ=0\gamma = 0. The attention and MLP output projections (the "writers" into the residual stream) are jittered together in one stream, the MLP's first layer in another, all at σ=0.2\sigma = 0.2.
  3. Attention outputs, block by block, with σ=0.2\sigma = 0.2 in the first half and 0.40.4 in the deeper half, because, in a code comment's words, "The deeper half moves the loss too little at 0.2 for a readable reward". These estimated errors become targets.
  4. Attention internals (queries, keys, values, the value-embedding gate and the value embeddings) never touch the loss. local_errors recomputes only that block's attention from cached activations with a jitter and scores the change against the estimated attention-output error, the post's cs=−⟨g^s,Δos⟩c_s = -\langle \hat g_s, \Delta o_s\rangle. Keys, values, gates and value embeddings sum the scores of later tokens decayed by γ=0.98\gamma = 0.98; queries keep their own token's score.
  5. The head is jittered on the cached logits, one 256-column slab of the 4,096-token vocabulary at a time, recomputing only the cross-entropy from cached partial sums.
  6. The 2L residual-mix scalars are the only weights trained by weight-space ES, with two-sided pairs.
  7. reassemble turns each error into a weight gradient with torch.einsum('btd,bti->di', error, x), or a scatter-add for an embedding. Every estimate is divided by the token count, written into .grad, and SGD with momentum takes the step.

So the gradient with respect to the weights is not inferred from weight noise at all. It is the outer product backprop forms and the animation shows, with only the error vector coming from the population.

The population count is narrower than it sounds. In recipe(), at population 256 the writer and MLP-hidden streams each get 110 draws spread over the eight blocks (38, 32, 16, 10, 6, 4, 2 and 2, tilted toward the first blocks), the attention outputs 26 and the token embedding 10, which sums to 256 (measured). At 16,384 everything scales by 64: 7,040 + 7,040 + 1,664 + 640 = 16,384, matching the appendix's Table 3 (measured). The head's 65,536 slab draws and the attention-internal draws are on top of KK; the post says they are a small share of an update's compute (reported).

One oddity: the token embedding trains at a learning rate of 1000 under SGD, for Dust and backprop alike (measured in make_optimizer, matching the appendix).

The results, checked against the tables

The setup (reported): FineWeb with a 4,096-token BPE tokenizer, batches of 8 sequences of 2,048 tokens, one epoch, SGD with momentum at a constant learning rate, three seeds per cell, every method tuned per token budget and population. In the released config, the 1M-token budget is 61 steps and the 10M budget 610 (measured).

Two line charts of validation loss against training fraction, at 1M tokens (left) and 10M tokens (right). Dust at populations 64, 1k and 16k in shades of violet; EGGROLL-Transformer at 64, 1k and 16k in shades of grey; backprop as a black dashed line. At 1M tokens Dust at 1k and 16k track backprop closely and end around 5.9, Dust at 64 ends near 6.3, and EGGROLL ends between about 6.7 and 7.4. At 10M tokens Dust at 1k and 16k hug the backprop curve down to about 5.0, Dust at 64 ends near 5.4, and EGGROLL flattens between about 6.0 and 7.1.
Validation loss during training at 1M and 10M tokens. Dust at 1k and 16k draws stays on backprop's curve; every EGGROLL curve, including 16k forward passes, flattens above Dust at 64 draws. Means over three seeds (Dust post, Figure 1).

The headline ladder is Table 1 in the appendix. Test loss at the best-validation checkpoint (reported):

BudgetDust, 64Dust, 1kDust, 16kBackpropEGGROLL, 16k
100k tokens7.2407.1517.1707.2167.263
1M tokens6.2905.9345.9165.9596.708
10M tokens5.3975.1115.0494.9896.033
20M tokens5.2654.9604.8024.6335.865

"Beats backprop at 100k and 1M tokens" holds in the table. The best Dust cell is 7.151 against 7.216 at 100k, and 5.911 (at 4k draws) against 5.959 at 1M, margins of 0.065 and 0.048 nats (reasoned). "Close at 10M and 20M" holds too, with gaps at 16k draws of 0.060 and 0.169 nats (reasoned). The 20M claim that the asymptote sits below backprop rests on a power-law fit through five cells that are still falling: limit 4.431, 95% interval 3.89 to 4.58, against backprop's 4.633 (reported). The post itself reads that as "evidence that the gap keeps closing with population rather than as a measured limit". So should you.

Four panels of test loss against population (64, 256, 1k, 4k, 16k) at 100k, 1M, 10M and 20M tokens. EGGROLL-Transformer in grey falls slowly and stays highest in every panel. Dust in violet falls quickly and flattens near the dotted backprop line, slightly below it at 100k and 1M, just above it at 10M, and above it at 20M, where a dash-dotted line marks the power-law fit's limit of 4.43 with a 95% interval bar below backprop.
Test loss against population at four token budgets. The dashed violet curve is a power-law fit through Dust's five cells; at 20M its limit, 4.43, is an extrapolation from a ladder that is still falling (Dust post, Figure 2).

The EGGROLL comparison is two different claims. The measured one is in the table: EGGROLL at 16k forward passes never reaches Dust at 64 draws at any budget. At 100k tokens it is 7.263 against 7.240, 0.023 behind, which the post rounds to "within 0.02"; at 1M, 10M and 20M it is 0.4 to 0.6 nats behind (reported). The "1,000 to 10,000x" figure is not measured; the post says "based on our extrapolations". Appendix Table 2 extends EGGROLL's ladder by four kinds of fit to the population it would need to match Dust at 64 draws. The main fits put it between 3,100x and 23,000x that population. Refitting with one population left out widens the range to 54,000x at 10M, and one such refit at 1M never gets there at all (reported). Q Labs' own caveat is that, from 256 draws up, a Dust draw costs less than an EGGROLL forward pass, so the comparison is lenient to EGGROLL (reported).

Under Adam, at 1M tokens only, the shape repeats. Dust's power-law limit is 5.248 with an interval of 5.17 to 5.30, below backprop; EGGROLL's Adam ladder lands within 0.01 of its SGD ladder at every population (reported). That is one budget, and the Adam runs give the token embedding extra draws, about 1.4x the listed population in total.

Bigger models need fewer draws, up to a point

The claim I found most interesting is the one about size. The variance argument above says a bigger model needs a bigger population. Q Labs trained four sizes at a fixed 10M tokens: 2.0M, 7.3M, 38M and 243M parameters (2, 4, 8 and 16 blocks at widths 128 to 1024).

A heat table of test loss with rows for populations 64, 256, 1k, 4k, 16k and backprop, and columns for 2.0M, 7.3M, 38M and 243M parameters, plus a sparkline per row. Bold marks the best size per row: 7.3M at population 64 (5.556), 38M at every other row (5.358, 5.158, 5.053, 5.036) and for backprop (5.015). The 243M column reads 5.719, 5.419, 5.214, 5.124, 5.086 and 5.048 for backprop. The 2.0M column reads 5.705, 5.486, 5.265, 5.189, 5.171 and 5.180 for backprop.
Four model sizes at a fixed 10M tokens, test loss against population; the best size at each population is bold. These runs used earlier training settings than the main comparison (Dust post, Figure 4).

Read straight off the table (measured from the published cells):

Q Labs' reading is that overparameterization gives search "a bigger space with better geometry". My reading is narrower. Per-token rewards make the search dimension per reward the layer's output width, which for square layers grows roughly like D\sqrt{D} rather than like DD, so the variance penalty for size is mild. A wider model is also a better model at a fixed token count, and that gain can outrun the penalty. One more caveat from the appendix: the size study predates the main protocol and uses the old token-embedding learning rate, chosen because at 1000 "backprop shows no clear separation between model sizes".

Backprop-like gradients, measured as a cosine

The last claim is that backprop-like gradients emerge from the population. Q Labs measured the cosine between Dust's estimate and the exact backprop gradient on the same batch, on backprop-trained checkpoints at 10M, 100M and 1B tokens, with populations from 64 to 128k.

Four panels of cosine to the backprop gradient against population from 64 to 128k. The first, EGGROLL at 10M tokens, has its own axis topping out near 0.12; only the head line rises above 0.05. The other three, Dust at 10M, 100M and 1B tokens, rise from near zero to between about 0.5 and 0.9 for the MLP, Q, K and V lines, and the head line rises earliest, saturating near 1.0 by 16k. A dotted vertical line marks the 16k population used in training.
Cosine between the estimated and backprop gradient per layer type, against population, for EGGROLL at 10M tokens (left, own axis) and Dust at three checkpoints. Dust's lines are fits of cos(K) = c_max / sqrt(1 + c/K) (Dust post, Figure 5).

A two-parameter law, cos⁡(K)=cmax⁡/1+c/K\cos(K) = c_{\max}/\sqrt{1 + c/K}, fits each layer type with RMSE under 0.06 (reported). It is the same K/(K+D)\sqrt{K/(K+D)} shape as the estimator variance above, with cc playing the effective dimension and cmax⁡c_{\max} a ceiling below one. Plugging the appendix's Table 8 fits for the 100M-token checkpoint into the law at the training population of 16,384 (reasoned): MLP output 0.51, attention output 0.57, token embedding 0.49, attention values 0.26, attention keys 0.16, and the head 0.98. The keys' fitted cc is 616,595, which is why their curve has barely left the floor at 16k. EGGROLL, measured the same way, stays below 0.05 at 128k for every layer type except the head (reported).

Two things to keep in mind. These are cosines on backprop's checkpoints, not along Dust's own trajectory. And a cosine of 0.5 is a modest alignment; SGD with momentum averages over steps, which is part of why it trains anyway. The post's further point, that Dust's gradients are "useful but different" and that the difference is sometimes better, is an observation from the ladders, not something the cosine study tests.

What it costs

This is where the result needs the most care. The post says it plainly: "We must have orders of magnitude more compute efficiency before Dust becomes a practical alternative to backprop at current levels of compute." Neither the post nor the code reports wall-clock time against backprop.

From the code I can estimate the forward work per update (reasoned). A draw at block ll reruns blocks ll through 8. With the 1M-token allocation at 16,384 draws, the writer and MLP-hidden streams, the attention outputs and the token embedding cost about 13,000 block-passes' worth, in units of one full eight-block forward. Backprop costs roughly three forwards' worth. That puts Dust at a few thousand times backprop's compute per update, before counting the head and attention-internal draws. At 256 draws the same arithmetic gives about 200 forward-equivalents, around 70x. The paper's eight-GPU runs split the draws across GPUs; they do not make them cheaper.

The regime is also early. The best losses here are 4.6 to 4.8 nats at 20M tokens, against a uniform-prediction loss of ln⁡4096≈8.32\ln 4096 \approx 8.32 (reasoned). One reply to the thread made the point that at 5 nats most of the signal is still pushing unlikely tokens down, and that feature learning has barely started. The tuned grid matters too: backprop here is SGD with momentum at a constant rate, not a tuned Muon or AdamW schedule. The Adam ladder helps, but only at 1M tokens.

Where it sits

Dust is honest about its lineage. Node perturbation dates to Widrow and Lehr; GEMINI (Le Cun, Galland and Hinton, 1988) injected noise into hidden layers; Ren et al. scaled forward gradients with local losses; MeZO fine-tunes with forward passes alone; and Zoop used output perturbations for fine-tuning; a reply to the thread linking that paper said they tried "something very similar last year but only saw it working on finetuning". Another reply pointed to Nesterov's random gradient-free methods and to a functional version for infinite-dimensional spaces (arXiv 2512.20566), from one of the authors of the functional gradient descent paper covered here. That work is a related estimator, not a source for Dust.

What is new is the combination: per-token rewards as a population axis, separate passes per layer type to keep interference down, attention internals credited through an estimated local target, and a tuning loop that maximises cosine to backprop on one batch before any training. The post's motivation is not efficiency. It is that a search-based optimizer does not need differentiability, so it could train looped or program-in-the-loop models that backpropagation through time struggles with. None of that is tested here. On a smaller scale, the same trade, evaluations instead of gradients, shows up in Sakana's Fugu, where CMA-ES tunes a coordinator over ten thousand parameters.

What I want next: a budget where backprop gets below 4 nats, wall-clock next to backprop on the same GPUs, and a model only Dust can train. Until then, Dust is the best evidence I have seen that zeroth-order credit assignment can follow backprop through transformer pretraining, at a price that is still its main problem.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Dust: every token is a member of the population", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026dustzerothorder,
  author = {Satyajit Ghana},
  title  = {Dust: every token is a member of the population},
  url    = {https://ai.thesatyajit.com/articles/dust-zeroth-order},
  year   = {2026}
}
share