~/satyajit

Olmo-core 3: an open training stack for trillion-parameter MoE

mdjsonmcp

2026-10-02 · 15 min · mixture-of-experts · systems · training · pretraining · open-source · llm · explainer

On the same day in October 2026, two groups published how they train large Mixture-of-Experts models. Ai2 released Olmo-core 3 — the training system behind the next generation of Olmo, with the code on GitHub (Apache-2.0) and a 168-page technical report. Aleph Alpha published a blog post walking through the parallelism choices they make as they scale a 30B-A3B MoE from 16 to 512 GPUs. Same problem, two very different answers — and two very different things handed to you to check.

They make a good pair because the thing they are both wrestling with is not a modeling problem. It is a systems problem. A big MoE is cheap in arithmetic and expensive in everything else, and the gap between those two is where a training stack lives or dies. This is a walk through what that gap is, the knobs you have to close it, and how far each team's numbers actually go. Every figure below is labeled measured (I computed it from the repo), reported (the publisher's number, not re-run), or reasoned (my arithmetic on the other two).

Why a big MoE is a systems problem

A Mixture-of-Experts layer replaces one feed-forward network with many — the experts — and a router that sends each token to only a few of them. The appeal is in the accounting: a model can carry a huge number of total parameters while each token only pays for the handful it activates. If you have not seen the mechanism built from first principles, my Mixture of Experts, from scratch walks through the gating, dispatch, and load-balancing; Switch Transformers is the top-1 simplification the modern MoE zoo descends from.

Here is the catch, and it is the first sentence of the Olmo-core 3 report's abstract: training cost depends not only on active compute but on "bytes moved, memory held, and kernels launched — and many of those costs follow total parameter capacity rather than active compute" (reported). Your FLOPs scale with the active parameters. Your weight storage, your gradient synchronization, and your optimizer updates scale with the total parameters. An MoE is designed to make those two numbers diverge. So the moment you add experts at a fixed active size — exactly what you do to make a model better without making each token more expensive — most of your costs go up and your useful arithmetic does not.

Ai2 calls this the capacity tax, and they measured it on the infrastructure they already had. Olmo-core 2 trained dense models with Fully Sharded Data Parallel (FSDP), which shards the model across GPUs and re-gathers each layer's weights just before it runs, then frees them. For a dense model that is a good trade. For an MoE it is a disaster: you re-gather the entire, growing expert pool around every microbatch, even though each token only touched a few experts.

Two line plots. Left: MFU rising with global batch size for DDP (25% to 41%) and FSDP (26% to 30%). Right: MFU versus total model size as the expert pool grows from 8 to 128 experts; MoE FSDP falls from 41% to 16%, MoE DDP without EP runs out of memory at 64 experts, and MoE DDP with EP8 stays flat near 42-44%.
The capacity tax, measured on 8 B300 GPUs with a fixed ~3.2B active parameters. As the expert pool grows from 8 to 128 (4.6B to 47B total params), MoE-FSDP's useful-model MFU collapses from 40.9% to 15.7% even though the work per token is unchanged; DDP+EP8 stays near 42-44% (Olmo-core 3 report, Figure 8).

Read the right panel. The pink line is MoE under FSDP: as the expert pool grows from 8 to 128 experts — total parameters from 4.6B to 47B — useful-model MFU falls from 40.9% to 15.7% (reported), while the active compute per token never changes. That is a 61.6% relative drop in efficiency bought with nothing (reasoned). The dense baselines (dashed, 49.8% for FSDP and 51.1% for DDP) sit far above it. The whole redesign exists to flatten that pink line.

The parallelism axes

To see how, you need the vocabulary — the handful of ways you can split a model across GPUs. They compose, and picking which to turn on is most of the job.

The one that is new to MoE is EP, and it is the hinge of Olmo-core 3's design. Instead of moving the weights to the data (FSDP's re-gather), you keep the experts resident and move the data to the weights. Toggle between the two:

8 experts · 4 GPUs · top-k routing
t1t2t3t4token rowsGPU 0E1E2GPU 1E3E4GPU 2E5E6GPU 3E7E8
Olmo-core 3 · Experts stay resident; EP sends only the selected token rows to the GPU that owns their expert (all-to-all dispatch), runs it, and sends results back (combine).

The difference is which thing moves. Olmo-core 2 moves the weights to the data, re-gathering the whole (growing) expert pool every microbatch. Olmo-core 3 keeps the experts put and moves the data to the weights — only the handful of token rows each expert was routed. That is the all-to-all dispatch and combine at the heart of expert parallelism, and it is why throughput stops falling as you add experts.

That all-to-all dispatch is the new communication pattern an MoE stack has to be built around. It is also the thing that broke when Ai2 tried to retrofit it onto dense-model plumbing, and the thing Aleph Alpha deliberately avoided. Same mechanism, opposite decisions — which is the whole reason the two writeups are worth reading together.

Olmo-core 3's answer: experts resident, data routed

Olmo-core 3 (v3.0.0, Apache-2.0 — measured, from the repo) throws out the full-reshard FSDP path and rebuilds on DDP. Each rank keeps its partition of the model resident for the whole gradient accumulation window and synchronizes gradients once per optimizer step. On top of that base it layers the axes that actually divide an MoE: EP to shard the routed experts, PP to split the layers when a rank cannot hold them, and a distributed optimizer that shards the FP32 main weights and optimizer moments. The thread's one-line version: Olmo-core 2 "repeatedly gathered model weights for each small batch"; Olmo-core 3 "keeps experts resident on GPUs and sends data to them" (reported).

The memory accounting is the cleanest way to see why that composition is the point.

Three stacked-byte diagrams. DDP alone: 18 bytes per parameter (2 BF16 model, 4 FP32 gradient, 4 FP32 main param, 4+4 optimizer moments). DDP plus distributed optimizer: 6 plus 12 over DP bytes. DDP plus EP plus distributed optimizer: 6 over EP-model-parallel plus 12 over DP bytes.
Each axis shaves a different part of the per-parameter memory bill. Naive DDP holds 18 bytes per parameter; sharding the FP32 main weights and optimizer states over D replica ranks drops it to 6 + 12/D; sharding the routed experts over the EP model-parallel degree takes their contribution to 6/M_E + 12/D (Olmo-core 3 report, Figure 9).

Under a BF16-compute, FP32-gradient, two-moment Adam policy, naive DDP costs 18 bytes per parameter (reported): 2 for the BF16 weight, 4 for the FP32 gradient, 4 for the FP32 main weight, and 8 for the two optimizer moments. The distributed optimizer shards the FP32 state over D replica ranks, so the bill becomes 6 + 12/D bytes — at D = 8 that is 7.5 bytes, a 2.4x reduction (reasoned). EP then shards the routed experts' contribution again, over the expert-parallel degree. Each axis divides a different part of the state, so what every GPU holds stays bounded as the global model grows. That is the sentence that makes trillion parameters feasible on 512 GPUs.

The engineering under those axes is where the report earns its length, and it is all in the repo if you want to read it rather than take it on faith:

allenai/olmo-core@a800c06 · snapshot 2026-10-02
tracked files
744
license
Apache-2.0
branch
main
tests
211 files
source
6.8 MB
commit date
2026-10-01
source by language
Python6.4 MB(586)CUDA231.2 kB(13)C++48.1 kB(8)Shell31.1 kB(14)Makefile14.9 kB(2)Dockerfile13.8 kB(1)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-02 at a800c06 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

The headline numbers, read with their conditions

A panel of five model scales named Tiny to Ultra, each showing active-at-total parameters, activation ratio, throughput in TFLOP/s per GPU, MFU percent, expert count, and the precision and EP/PP topology. From Tiny 1.59B at 12.9B (903 TFLOP/s, 40.1% MFU) to Ultra 58.4B at 1.2T (858 TFLOP/s, 38.1% MFU).
Selected achieved throughput on B300 NVL8, from Tiny (12.9B total) on 16 GPUs to Ultra (1.2T total) on 512 GPUs. MFU is normalized to the B300 dense BF16 peak of 2,250 TFLOP/s/GPU — a BF16 reference even for the MXFP8 row, not FP8-peak utilization (Olmo-core 3 report, Table 30).

The report's headline is a five-point matrix (all reported): Tiny (1.59B active at 12.9B total) hits 903 TFLOP/s/GPU and 40.1% MFU on 16 GPUs with plain DDP; Ultra (58.36B active at 1.2T total) reaches 858 TFLOP/s/GPU and 38.1% MFU on 512 GPUs with MXFP8, EP8, PP8, and per-layer recompute. In between, throughput barely moves as total capacity grows by two orders of magnitude — which is the capacity tax being paid down. For the same model on the same GPUs, Olmo-core 3 delivers about 2.7x the throughput of Olmo-core 2 (reported, blog and thread); the blog pins one point of that to a 47B MoE going from 19,400 to 52,000 tokens per second per GPU on 8 B300 (reported).

Two honesties are baked into how they present this, and they matter. First, every headline rate uses random routing as a systems reference — it removes a learned router's shifting expert loads to give a stable number — so these are not time-to-quality claims, and the report says the learned-router companions run about 1% slower at Tiny and 9% slower at Small (reported). Second, the five configurations differ in batch, precision, and recompute, so the matrix is a set of feasible operating points, "not a controlled scaling curve." An optional DeepEP v2 backend reaches a 2.38T-parameter configuration, but from a nine-step run "whose curve had not stabilized" — reachable capacity, not sustained throughput (reported). I appreciate a report that labels its own numbers more conservatively than its press would.

What MFU is, and why 35% is worth stating

Both teams lead with Model FLOPs Utilization, so it is worth being precise about what it measures:

MFU=useful model FLOP/s delivered per GPUhardware peak FLOP/s per GPU\text{MFU} = \frac{\text{useful model FLOP/s delivered per GPU}}{\text{hardware peak FLOP/s per GPU}}

"Useful model FLOPs" are the arithmetic the model definition requires for a forward and backward pass — not the FLOPs you actually executed, which recompute and padding inflate. The denominator is the chip's advertised peak. Ai2 uses the B300 dense BF16 peak of 2,250 TFLOP/s/GPU, and keeps that same denominator for the MXFP8 row, so 858 / 2250 = 38.1% is deliberately a BF16 reference, not FP8-peak utilization (reasoned, matching the reported value).

Why care? Because MFU is the one number that resists hand-waving. "Blazing fast" tells you nothing; "38% of a 2,250 TFLOP/s/GPU peak, denominator stated, on 512 GPUs" tells you exactly how much silicon is doing useful work. Frontier dense training typically lands somewhere in the 30-50% range; MoE makes it harder because the all-to-all dispatch and the capacity tax eat into it. A measured, denominator- disclosed 35% at 512 GPUs is a real result — and most labs never publish the number at all.

Aleph Alpha's recipe: the opposite choice, and why it works

Which is what makes Aleph Alpha's writeup the right foil. They take a 30B-A3B MoE (30B total, 3B active — reported) and scale it across NVIDIA B200 GPUs, 8 per node over NVLink, nodes joined by InfiniBand, from 16 to 512 GPUs. And they turn on almost nothing. No TP, no EP, no PP, no CP — just FSDP plus data parallel, with selective activation checkpointing throughout.

parallelism planner · reported operating points
MoE, up to 1.2T · NVIDIA B300 (NVL8)Tiny · 1.59B@12.9B · 16 GPUs
DP
replica · split the batch
FSDP
shard + re-gather weights
TP
split weight matrices
EP
experts across GPUs · all-to-all
PP
layers into pipeline stages
CP
split the sequence
MFU
40.1%
useful-model throughput
903TFLOP/s/GPU
MFU across this campaign’s measured points
GPUs →

64 experts, BF16, 16 GPUs. Small enough that plain DDP is all you need — no EP, no PP.

Reported operating points, not a controlled scaling curve: the configs differ in batch, precision and recompute. Aleph Alpha’s points are a 16–512-GPU sweep of one model; Olmo-core’s five sizes each run on ≤512 GPUs. Numbers from the Aleph Alpha blog and the Olmo-core 3 report.

Slide through their three operating points. At 16 GPUs, FSDP shards the whole model across all 16; at 128, FSDP stretches to 128 ranks, which they diagnose from profiler traces as "right at the boundary of being communication-bound" (reported); at 512, they cap FSDP at 128 and replicate it four ways with plain DP — hybrid sharding, HSDP — because DP's extra cost is "paid only in the last gradient accumulation step" and so stays cheap. The result: MFU of 37.5% / 37.0% / 35.3% and per-GPU throughput of 28.4k / 28.0k / 26.7k tokens/second across 16 / 128 / 512 GPUs (reported).

Per-GPU throughput falling from 28.4k to 26.7k across a 32x scale-up is a 6.0% drop (reasoned), which is their "only 6% below perfect linear scaling" — the aggregate box in the planner shows it as 94% of a linear line. That is what "near-linear" has to mean to be a real claim: you added 32x the GPUs and kept 94% of the per-GPU rate. Their projected run — 20 trillion tokens on 512 GPUs in 17 days — cross-checks cleanly: 20e12 tokens over 17 days across 512 GPUs is 26,600 tokens/second/GPU (reasoned), right on their reported 26.7k.

The reason they can skip EP is the single most useful sentence in their post: "If our model were larger or sparser, we would likely need EP, as the communication-to-computation ratio would not tilt in our favour" (reported). A 30B-A3B model is small and dense enough that FSDP's weight re-gather is not yet the bottleneck Ai2 measured at 47B and beyond. That is the whole lesson, stated as a threshold: the parallelism you need is a function of model size and sparsity against your hardware's interconnect, not a fixed recipe. Flip the planner between the two campaigns and you are watching that function evaluate at two different inputs — Aleph Alpha lighting FSDP, Ai2 lighting DP + EP + PP.

The two meanings of open

There is a last axis these two share, and it is the one the site cares about most: both are open, but not in the same way, and the difference is instructive.

Olmo-core is open as code. Apache-2.0, on GitHub, with a 168-page report that documents not only what worked but what did not — and the failures are the most valuable part. They name a routing pathology they call Token Gerrymandering, where the router learns to lower the load-balancing loss while making the actual expert imbalance worse (reported). They report that overlapping communication with computation can slow the overlapped kernels "by more than it hides" (reported) — a negative result that will save someone a week. You can clone it, read the kernel, and run it on your own hardware. This is the same open-infrastructure instinct behind MegaTrain's single-GPU training and the stack that made Ring-Zero's trillion-parameter RL stable; the MoE-RL failure mode in Rollout Routing Replay is a cousin of Token Gerrymandering, a router that behaves differently than you think it does.

Aleph Alpha is open as recipe. No code, but something code alone does not give you: the reasoning. They show the PyTorch profiler traces at each scale, point at the exact kernels, and explain why they cap FSDP at 128 and add DP. You cannot run it, but you can learn the decision procedure — which, for a practitioner choosing their own parallelism, may be worth as much as a repo.

Neither is the closed-lab norm of "we trained a big MoE, here is the model." Both tell you how. If you want the full openness ledger for an MoE — code, checkpoints, data recipes, the lot — AMD's Instella-MoE is the maximal version; and if you care about where the scaling budget itself is headed, scaling laws in 2026 is why these teams are building for the trillion-parameter range at all.

What an open MoE training stack has to get right

Strip it to the load-bearing ideas:

  1. Stop paying the capacity tax. Keep experts resident and route tokens to them (EP's all-to-all), rather than re-gathering a growing expert pool every microbatch. This is what flattens the pink line in Figure 8, and it is the difference between Olmo-core 2 and 3.
  2. Compose axes that divide different state. DDP for the base, EP for the experts, PP for the layers, a distributed optimizer for the FP32 state — so per-GPU memory stays bounded as the global model grows. 6/M_E + 12/D bytes per parameter is the formula that reaches 1.2T on 512 GPUs.
  3. Report MFU with its denominator, and near-linear with its reference. 38.1% of a stated 2,250 TFLOP/s/GPU peak; 94% of a per-GPU-flat linear line. Numbers you can argue with beat adjectives you cannot.
  4. Match the parallelism to the model. A 30B-A3B gets to 35% MFU near-linearly with only FSDP+DP; a 1.2T MoE needs EP+PP+MXFP8. Neither recipe is wrong; they are the same function at different inputs.

The quiet thing both releases demonstrate is that the hard part of large-model training has moved from the model to the plumbing — the bytes moved, the kernels launched, the host waits removed. Olmo-core 3 hands you the plumbing and a report on how it bends. Aleph Alpha hands you the reasoning for bending it a different way. Read together, they are about as honest a picture of open MoE training as you will get without an H-cluster of your own.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Olmo-core 3: an open training stack for trillion-parameter MoE", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026olmocore3,
  author = {Satyajit Ghana},
  title  = {Olmo-core 3: an open training stack for trillion-parameter MoE},
  url    = {https://ai.thesatyajit.com/articles/olmo-core-3},
  year   = {2026}
}
share