2026-10-02 · 15 min · mixture-of-experts · systems · training · pretraining · open-source · llm · explainer
On the same day in October 2026, two groups published how they train large Mixture-of-Experts models. Ai2 released Olmo-core 3 — the training system behind the next generation of Olmo, with the code on GitHub (Apache-2.0) and a 168-page technical report. Aleph Alpha published a blog post walking through the parallelism choices they make as they scale a 30B-A3B MoE from 16 to 512 GPUs. Same problem, two very different answers — and two very different things handed to you to check.
They make a good pair because the thing they are both wrestling with is not a modeling problem. It is a systems problem. A big MoE is cheap in arithmetic and expensive in everything else, and the gap between those two is where a training stack lives or dies. This is a walk through what that gap is, the knobs you have to close it, and how far each team's numbers actually go. Every figure below is labeled measured (I computed it from the repo), reported (the publisher's number, not re-run), or reasoned (my arithmetic on the other two).
Why a big MoE is a systems problem
A Mixture-of-Experts layer replaces one feed-forward network with many — the experts — and a router that sends each token to only a few of them. The appeal is in the accounting: a model can carry a huge number of total parameters while each token only pays for the handful it activates. If you have not seen the mechanism built from first principles, my Mixture of Experts, from scratch walks through the gating, dispatch, and load-balancing; Switch Transformers is the top-1 simplification the modern MoE zoo descends from.
Here is the catch, and it is the first sentence of the Olmo-core 3 report's abstract: training cost depends not only on active compute but on "bytes moved, memory held, and kernels launched — and many of those costs follow total parameter capacity rather than active compute" (reported). Your FLOPs scale with the active parameters. Your weight storage, your gradient synchronization, and your optimizer updates scale with the total parameters. An MoE is designed to make those two numbers diverge. So the moment you add experts at a fixed active size — exactly what you do to make a model better without making each token more expensive — most of your costs go up and your useful arithmetic does not.
Ai2 calls this the capacity tax, and they measured it on the infrastructure they already had. Olmo-core 2 trained dense models with Fully Sharded Data Parallel (FSDP), which shards the model across GPUs and re-gathers each layer's weights just before it runs, then frees them. For a dense model that is a good trade. For an MoE it is a disaster: you re-gather the entire, growing expert pool around every microbatch, even though each token only touched a few experts.

Read the right panel. The pink line is MoE under FSDP: as the expert pool grows from 8 to 128 experts — total parameters from 4.6B to 47B — useful-model MFU falls from 40.9% to 15.7% (reported), while the active compute per token never changes. That is a 61.6% relative drop in efficiency bought with nothing (reasoned). The dense baselines (dashed, 49.8% for FSDP and 51.1% for DDP) sit far above it. The whole redesign exists to flatten that pink line.
The parallelism axes
To see how, you need the vocabulary — the handful of ways you can split a model across GPUs. They compose, and picking which to turn on is most of the job.
- DP (data parallel). Replicate the whole model on each GPU, split the batch, average the gradients with an all-reduce once per step. Simple, communication-light, but every GPU must hold the whole model.
- FSDP (fully sharded data parallel, a.k.a. ZeRO-3). Shard the parameters, gradients, and optimizer state across the DP group; all-gather each layer's weights just-in-time, use them, free them. Trades memory for a weight all-gather on every microbatch.
- TP (tensor parallel). Split the individual weight matrices across GPUs and all-reduce inside each layer. Heavy, latency-sensitive communication; kept inside one node.
- EP (expert parallel). The MoE-native axis. Put different experts on different GPUs and send each token to the GPU that owns its expert — an all-to-all dispatch, then an all-to-all combine to bring the results back.
- PP (pipeline parallel). Split the layers into stages across GPUs and flow microbatches through the pipeline. Needed once one GPU cannot hold its share of the layers; costs a pipeline "bubble".
- CP (context parallel). Split the sequence dimension so a single long example spans GPUs. For long context, not for capacity.
The one that is new to MoE is EP, and it is the hinge of Olmo-core 3's design. Instead of moving the weights to the data (FSDP's re-gather), you keep the experts resident and move the data to the weights. Toggle between the two:
The difference is which thing moves. Olmo-core 2 moves the weights to the data, re-gathering the whole (growing) expert pool every microbatch. Olmo-core 3 keeps the experts put and moves the data to the weights — only the handful of token rows each expert was routed. That is the all-to-all dispatch and combine at the heart of expert parallelism, and it is why throughput stops falling as you add experts.
That all-to-all dispatch is the new communication pattern an MoE stack has to be built around. It is also the thing that broke when Ai2 tried to retrofit it onto dense-model plumbing, and the thing Aleph Alpha deliberately avoided. Same mechanism, opposite decisions — which is the whole reason the two writeups are worth reading together.
Olmo-core 3's answer: experts resident, data routed
Olmo-core 3 (v3.0.0, Apache-2.0 — measured, from the repo) throws out the full-reshard FSDP path
and rebuilds on DDP. Each rank keeps its partition of the model resident for the whole gradient
accumulation window and synchronizes gradients once per optimizer step. On top of that base it
layers the axes that actually divide an MoE: EP to shard the routed experts, PP to split the
layers when a rank cannot hold them, and a distributed optimizer that shards the FP32 main
weights and optimizer moments. The thread's one-line version: Olmo-core 2 "repeatedly gathered model
weights for each small batch"; Olmo-core 3 "keeps experts resident on GPUs and sends data to them"
(reported).
The memory accounting is the cleanest way to see why that composition is the point.

Under a BF16-compute, FP32-gradient, two-moment Adam policy, naive DDP costs 18 bytes per
parameter (reported): 2 for the BF16 weight, 4 for the FP32 gradient, 4 for the FP32 main weight,
and 8 for the two optimizer moments. The distributed optimizer shards the FP32 state over D replica
ranks, so the bill becomes 6 + 12/D bytes — at D = 8 that is 7.5 bytes, a 2.4x reduction
(reasoned). EP then shards the routed experts' contribution again, over the expert-parallel degree.
Each axis divides a different part of the state, so what every GPU holds stays bounded as the global
model grows. That is the sentence that makes trillion parameters feasible on 512 GPUs.
The engineering under those axes is where the report earns its length, and it is all in the repo if you want to read it rather than take it on faith:
- NVSHMEM rowwise expert parallel. The dispatch writes each token row directly into its
destination expert's buffer with GPU-initiated one-sided communication, avoiding the host-visible
split lists a block all-to-all needs (
src/olmo_core/nn/moe/v2/ep_no_sync_rowwise.py,comm.py, and theolmo_symm_mem_all_to_all.cuhkernel — measured). - Synchronization-free execution. The original MoE path copied routing counts to the CPU for
all-to-all split sizes and grouped-GEMM group sizes. Keeping that metadata on the device, and using
device-scheduled grouped GEMMs, removes two recurring host waits from the steady-state loop — the
ep_no_sync_*files are named for exactly this. - Device-scheduled grouped GEMM runs the uneven per-expert shapes routing produces without padding every expert's batch to a common capacity.
- MXFP8 — an 8-bit format with shared block scales — cuts GEMM time, saved-activation memory, and dispatch payloads, while one FP32 main weight stays authoritative.
- Topology-agnostic checkpoints store the global FP32 tensors independently of the parallel layout, so a run can resume on a different number of GPUs with a different EP/PP split.
- license
- Apache-2.0
- branch
- main
- tests
- 211 files
- source
- 6.8 MB
- commit date
- 2026-10-01
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-02 at a800c06 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
The headline numbers, read with their conditions

The report's headline is a five-point matrix (all reported): Tiny (1.59B active at 12.9B total) hits 903 TFLOP/s/GPU and 40.1% MFU on 16 GPUs with plain DDP; Ultra (58.36B active at 1.2T total) reaches 858 TFLOP/s/GPU and 38.1% MFU on 512 GPUs with MXFP8, EP8, PP8, and per-layer recompute. In between, throughput barely moves as total capacity grows by two orders of magnitude — which is the capacity tax being paid down. For the same model on the same GPUs, Olmo-core 3 delivers about 2.7x the throughput of Olmo-core 2 (reported, blog and thread); the blog pins one point of that to a 47B MoE going from 19,400 to 52,000 tokens per second per GPU on 8 B300 (reported).
Two honesties are baked into how they present this, and they matter. First, every headline rate uses random routing as a systems reference — it removes a learned router's shifting expert loads to give a stable number — so these are not time-to-quality claims, and the report says the learned-router companions run about 1% slower at Tiny and 9% slower at Small (reported). Second, the five configurations differ in batch, precision, and recompute, so the matrix is a set of feasible operating points, "not a controlled scaling curve." An optional DeepEP v2 backend reaches a 2.38T-parameter configuration, but from a nine-step run "whose curve had not stabilized" — reachable capacity, not sustained throughput (reported). I appreciate a report that labels its own numbers more conservatively than its press would.
What MFU is, and why 35% is worth stating
Both teams lead with Model FLOPs Utilization, so it is worth being precise about what it measures:
"Useful model FLOPs" are the arithmetic the model definition requires for a forward and backward pass — not the FLOPs you actually executed, which recompute and padding inflate. The denominator is the chip's advertised peak. Ai2 uses the B300 dense BF16 peak of 2,250 TFLOP/s/GPU, and keeps that same denominator for the MXFP8 row, so 858 / 2250 = 38.1% is deliberately a BF16 reference, not FP8-peak utilization (reasoned, matching the reported value).
Why care? Because MFU is the one number that resists hand-waving. "Blazing fast" tells you nothing; "38% of a 2,250 TFLOP/s/GPU peak, denominator stated, on 512 GPUs" tells you exactly how much silicon is doing useful work. Frontier dense training typically lands somewhere in the 30-50% range; MoE makes it harder because the all-to-all dispatch and the capacity tax eat into it. A measured, denominator- disclosed 35% at 512 GPUs is a real result — and most labs never publish the number at all.
Aleph Alpha's recipe: the opposite choice, and why it works
Which is what makes Aleph Alpha's writeup the right foil. They take a 30B-A3B MoE (30B total, 3B active — reported) and scale it across NVIDIA B200 GPUs, 8 per node over NVLink, nodes joined by InfiniBand, from 16 to 512 GPUs. And they turn on almost nothing. No TP, no EP, no PP, no CP — just FSDP plus data parallel, with selective activation checkpointing throughout.
64 experts, BF16, 16 GPUs. Small enough that plain DDP is all you need — no EP, no PP.
Slide through their three operating points. At 16 GPUs, FSDP shards the whole model across all 16; at 128, FSDP stretches to 128 ranks, which they diagnose from profiler traces as "right at the boundary of being communication-bound" (reported); at 512, they cap FSDP at 128 and replicate it four ways with plain DP — hybrid sharding, HSDP — because DP's extra cost is "paid only in the last gradient accumulation step" and so stays cheap. The result: MFU of 37.5% / 37.0% / 35.3% and per-GPU throughput of 28.4k / 28.0k / 26.7k tokens/second across 16 / 128 / 512 GPUs (reported).
Per-GPU throughput falling from 28.4k to 26.7k across a 32x scale-up is a 6.0% drop (reasoned), which is their "only 6% below perfect linear scaling" — the aggregate box in the planner shows it as 94% of a linear line. That is what "near-linear" has to mean to be a real claim: you added 32x the GPUs and kept 94% of the per-GPU rate. Their projected run — 20 trillion tokens on 512 GPUs in 17 days — cross-checks cleanly: 20e12 tokens over 17 days across 512 GPUs is 26,600 tokens/second/GPU (reasoned), right on their reported 26.7k.
The reason they can skip EP is the single most useful sentence in their post: "If our model were larger or sparser, we would likely need EP, as the communication-to-computation ratio would not tilt in our favour" (reported). A 30B-A3B model is small and dense enough that FSDP's weight re-gather is not yet the bottleneck Ai2 measured at 47B and beyond. That is the whole lesson, stated as a threshold: the parallelism you need is a function of model size and sparsity against your hardware's interconnect, not a fixed recipe. Flip the planner between the two campaigns and you are watching that function evaluate at two different inputs — Aleph Alpha lighting FSDP, Ai2 lighting DP + EP + PP.
The two meanings of open
There is a last axis these two share, and it is the one the site cares about most: both are open, but not in the same way, and the difference is instructive.
Olmo-core is open as code. Apache-2.0, on GitHub, with a 168-page report that documents not only what worked but what did not — and the failures are the most valuable part. They name a routing pathology they call Token Gerrymandering, where the router learns to lower the load-balancing loss while making the actual expert imbalance worse (reported). They report that overlapping communication with computation can slow the overlapped kernels "by more than it hides" (reported) — a negative result that will save someone a week. You can clone it, read the kernel, and run it on your own hardware. This is the same open-infrastructure instinct behind MegaTrain's single-GPU training and the stack that made Ring-Zero's trillion-parameter RL stable; the MoE-RL failure mode in Rollout Routing Replay is a cousin of Token Gerrymandering, a router that behaves differently than you think it does.
Aleph Alpha is open as recipe. No code, but something code alone does not give you: the reasoning. They show the PyTorch profiler traces at each scale, point at the exact kernels, and explain why they cap FSDP at 128 and add DP. You cannot run it, but you can learn the decision procedure — which, for a practitioner choosing their own parallelism, may be worth as much as a repo.
Neither is the closed-lab norm of "we trained a big MoE, here is the model." Both tell you how. If you want the full openness ledger for an MoE — code, checkpoints, data recipes, the lot — AMD's Instella-MoE is the maximal version; and if you care about where the scaling budget itself is headed, scaling laws in 2026 is why these teams are building for the trillion-parameter range at all.
What an open MoE training stack has to get right
Strip it to the load-bearing ideas:
- Stop paying the capacity tax. Keep experts resident and route tokens to them (EP's all-to-all), rather than re-gathering a growing expert pool every microbatch. This is what flattens the pink line in Figure 8, and it is the difference between Olmo-core 2 and 3.
- Compose axes that divide different state. DDP for the base, EP for the experts, PP for the
layers, a distributed optimizer for the FP32 state — so per-GPU memory stays bounded as the global
model grows.
6/M_E + 12/Dbytes per parameter is the formula that reaches 1.2T on 512 GPUs. - Report MFU with its denominator, and near-linear with its reference. 38.1% of a stated 2,250 TFLOP/s/GPU peak; 94% of a per-GPU-flat linear line. Numbers you can argue with beat adjectives you cannot.
- Match the parallelism to the model. A 30B-A3B gets to 35% MFU near-linearly with only FSDP+DP; a 1.2T MoE needs EP+PP+MXFP8. Neither recipe is wrong; they are the same function at different inputs.
The quiet thing both releases demonstrate is that the hard part of large-model training has moved from the model to the plumbing — the bytes moved, the kernels launched, the host waits removed. Olmo-core 3 hands you the plumbing and a report on how it bends. Aleph Alpha hands you the reasoning for bending it a different way. Read together, they are about as honest a picture of open MoE training as you will get without an H-cluster of your own.