~/satyajit

Reflection Beam: a 501B open MoE that spends 23B per token, and what its efficiency chart measures

mdjsonmcp

2026-10-06 · 18 min · llm · mixture-of-experts · reinforcement-learning · agentic-coding · open-weights · scaling · explainer

Reflection AI announced Beam on 5 October 2026, with a thread on X that opens: "a highly efficient agentic open model with 501B total parameters and 23B active." It is Reflection's first open-weight model. The pitch has three parts: it was trained end-to-end from scratch, it is "3-4x more efficient than GLM 5.2", and its RL run is "the largest publicly documented RL run we're aware of", 10.5k GB300s for four weeks.

The first thing to say is what is not there. As of today, 6 October, there are no weights. The Hugging Face API returns an empty list for reflection-ai, ReflectionAI and every spelling I tried. The blog says the weights, "technical report, model card, and developer artifacts" ship "later this month" under Apache 2.0, with FP8 and NVFP4 quantizations; the model is "undergoing final red-teaming", and early access is a waitlist at platform.reflection.ai. So this is a launch announcement with charts, not a release. Nothing below can be checked against a config file or a safetensors header, because neither exists in public yet.

What can be checked is the internal arithmetic of the announcement, and it turns out there is a lot of it. The blog publishes enough axis labels, footnotes and hardware counts to rebuild its efficiency chart from first principles and to put the training run on a ledger.

Labels, as everywhere on this site. Reported is Reflection's figure, not re-run. Measured is something I computed from a file: here, pixel positions read off Reflection's own published charts. Reasoned is my arithmetic on the other two.

Six grouped bar charts comparing Beam with Qwen 3.8 Max, GLM 5.2, Inkling and Nemotron Ultra on DeepSWE v1.1, Terminal Bench v2.1, HLE no tools, SWE Bench Pro V1, SWE Bench Verified and CritPT AA. Beam is highlighted in yellow-green; Chinese open models are grey and Western open models dark green.
The launch image: Beam against two Chinese and two Western open models on six benchmarks. Beam leads only SWE Bench Verified (80.9), where GLM 5.2 and Qwen 3.8 Max are 'not reported'. All scores are Reflection's, not re-run (Reflection, launch thread on X, image 1).

What is disclosed about the architecture

Not much in numbers, quite a lot in recipe. From the blog's "Stable and balanced MoE optimization" section (all reported):

That is enough to know the shape of the thing and not enough to compute anything about it: no hidden size, no head counts, no vocabulary, so no KV-cache arithmetic and no parameter recount of the kind I did for Kolibri-1. That waits for config.json.

Two of the stability claims come with plots, and they are worth looking at because they are the kind of evidence that usually stays internal.

A line chart of the busiest expert's load divided by uniform load, averaged across MoE layers, over training progress. It jumps to about 1.9x early, stays near there until about 20% of training, then declines steadily to 1.04x at the end.
Busiest-expert load relative to perfectly uniform routing, averaged over MoE layers. It peaks near 2x early and ends at 1.04x, the 'almost-perfect uniform utilization' the blog claims (Reflection, Beam blog, Figure 7).
Two panels of residual-stream RMS per layer on a log scale over pretraining, one after attention and one after MoE. Fifty-two lines, layer 1 lowest near 0.05 and layer 52 highest near 2, rise early and then flatten without growing.
Residual-stream RMS for all 52 layers, after attention and after the MoE block. Deeper layers sit higher, but no layer keeps growing, which is what the depth scaling and FP32 accumulation are for (Reflection, Beam blog, Figure 8).

What "efficiency" means here

The headline claim is from the thread: "3-4x more efficient than GLM 5.2 and over 4x more efficient than leading Western open models." The question to ask of any efficiency claim is: per what?

The blog answers it in the caption of its Figure 2, which is unusually explicit (reported):

estimating generation forward-pass compute as FLOPs ≈ 2 × active parameter count × mean generated tokens per attempt ... These estimates exclude prompt prefill, context-dependent attention operations, and serving overhead

So the x-axis is not measured cost, wall-clock or dollars. It is a formula:

F≈2⋅Nactive⋅TgenF \approx 2 \cdot N_{\text{active}} \cdot T_{\text{gen}}

where NactiveN_{\text{active}} is parameters touched per token and TgenT_{\text{gen}} is the mean number of tokens the model generates per attempt, reasoning included. The factor 2 is one multiply and one add per weight per token. The scores for the other models come from Artificial Analysis and DataCurve.

That formula has two inputs, and only one of them is behaviour. To see how much each contributes, I read the dots off both versions of the chart Reflection published: one with a token axis, one with the FLOP axis.

Three scatter panels of score against mean generated tokens for DeepSWE v1.1, HLE text-only and Terminal Bench 2.1. Beam's effort levels form a dark curve on the left of each panel; GLM-5.2 and Qwen3.8 sit further right.
The token version of the efficiency chart. Each Beam dot is a reasoning-effort setting. On Terminal Bench 2.1 Beam's top setting uses about 23k tokens; GLM-5.2's point sits near 81k (Reflection, Beam blog, 'Reasoning effort: score against mean generated tokens').
The same three panels with the x-axis changed to estimated generation forward FLOPs per attempt in PFLOP. Beam's curves are compressed against the left edge; GLM-5.2 sits around 3 to 6.5 PFLOP and Qwen3.8 further right.
The FLOP version, Reflection's Figure 2. Same dots, x-axis multiplied by 2 x active parameters, which is why the gap widens (Reflection, Beam blog, Figure 2).

I located each dot's pixel centre and mapped it through the gridlines (measured; good to about half a point and half a thousand tokens). Then I recomputed the FLOP axis myself from the token readings, with 23B for Beam and 40B for GLM-5.2. My recomputed FLOPs land within about 3% of where Reflection drew them: Beam's top DeepSWE point at 1.98 PFLOP against 2.04 read off their chart, GLM-5.2 at 6.22 against 6.30, GLM-5.2 on Terminal Bench at 6.48 against 6.54 (measured). The same back-solve gives 41B for Inkling's Terminal Bench point and 95B for Qwen3.8's HLE point, the active counts this site reports for Inkling and Qwen3.8-Max. The chart is the formula, applied consistently. Nothing hidden.

Now split the ratio. Any FLOP ratio against GLM-5.2 is the token ratio times 40/23=1.7440 / 23 = 1.74 (reasoned). That 1.74 is fixed by architecture: it is what you get from having fewer active parameters, before the model does anything clever. The rest is how many tokens each model spends:

BenchmarkGLM-5.2 (top point)Beam (comparison point)Token ratioFLOP ratio
DeepSWE v1.143.7% at 77.8k tokens44.3% at 43.1k1.8x3.1x
Terminal Bench 2.177.6% at 81.0k79.4% at 23.5k3.4x6.0x
HLE (text-only)41.0% at 40.7k35.7% at 19.4k2.1x3.7x

Scores and tokens are measured off the chart; the ratios are reasoned. On DeepSWE and Terminal Bench Beam's comparison point scores at least as high as GLM-5.2's, so those are fair like-for-like ratios. On HLE Beam never reaches GLM-5.2's score at any effort level; the 3.7x there buys GLM-5.2 about five more points, which is not the same thing as being 3.7x less efficient.

So "3-4x" is a fair summary of DeepSWE (3.1x) and conservative for Terminal Bench (6.0x), as long as you remember the unit is the formula, and that on DeepSWE about half the ratio, on a log scale, is the 1.74x active-parameter factor rather than shorter reasoning.

Beam vs GLM-5.2, rebuilt from Reflection's points
x-axis:
60%65%70%75%80%85%01234567estimated generation FLOPs per attempt (PFLOP)
Beam: 79.4% at 1.08 PFLOPGLM-5.2: 77.6% at 6.48 PFLOPGLM spends 3.4x the tokens, 6.0x the FLOPsa matched comparison: Beam scores at least as high here

The widget plots my readings of the dots. Switch the axis to tokens and the gap shrinks to what the model's behaviour alone accounts for; switch back and every GLM-5.2 point slides right by a factor of 40/23 relative to Beam's. Drag Beam's effort down and the comparison stops being matched: Beam at its lowest Terminal Bench setting is at 64.0% on about 7k tokens.

What the formula leaves out

The caption says it plainly and it matters for agentic work: prefill and attention are excluded. A Terminal Bench or SWE task is mostly prefill. The agent re-reads a growing transcript of file contents and command output on every turn, and attention cost grows with that context. A metric that counts only generated tokens rewards a model for being terse; it says nothing about the cost of the context it reads. With prefix caching that cost shrinks, but not to zero.

It also leaves out memory. The FLOP count uses 23B, but serving Beam means holding all 501B in memory, about 501 GB at FP8 or roughly half that at NVFP4 (reasoned, one byte or half a byte per weight, ignoring scales). GLM-5.2 holds 744B. On a per-node basis Beam's smaller total is a real advantage too, but it is not the advantage the chart measures.

None of this makes the claim wrong. It makes it a claim about tokens per task, scaled by active size, which is a useful thing to know and not the same as "completes tasks faster and cheaper", the thread's wording. That depends on the serving stack, which nobody outside Reflection can run until the weights ship.

Where Beam is and is not ahead

The blog's full table carries more models than the launch image, and it is candid in places. On Terminal Bench v2.1 Beam scores 80.1, against 81.0 for GLM-5.2, 86.6 for Qwen 3.8 Max, 88.3 for Kimi K3 and 90.6 for DeepSeek V4.1 Flash. On DeepSWE v1.1 it is 44.4, against 74.2 for DeepSeek V4.1 Flash. The blog says as much: "Where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time." (All reported.)

Two consistency notes. The launch image gives Beam 15.6 on CritPT; the blog's table says 16.3. And the GLM-5.2 Terminal Bench dot on the efficiency chart sits near 77.6%, while the table lists 81.0; the chart's points for other models come from Artificial Analysis, the table's presumably from the model's own report. Neither changes the story. Both are the kind of thing a technical report should settle.

The comparison set is chosen with care, too. "Over 4x more efficient than leading Western open models" compares Beam with Inkling, Nemotron 3 Ultra and Muse Glimmer, which on Terminal Bench also score 16 to 28 points lower (reported for the first two, 63.8 and 56.4 against 80.1; measured off the chart for Muse Glimmer, about 52%). That is a frontier claim more than an efficiency one, and the Pareto picture says it fairly: Beam is up and to the left of all three.

The RL run

The thread's strongest claim is about scale, and here the blog gives numbers I can do arithmetic on (all reported unless marked):

Some things follow directly. 100 million rollouts in 28 days is about 41 rollouts finishing every second. With 110K rollouts in flight on average, Little's law gives a mean rollout lifetime of about 2,660 seconds, roughly 44 minutes (reasoned, an upper bound since "over 100 million"). That is what long-horizon agentic RL looks like at the systems level: most of the fleet is waiting on rollouts that take most of an hour. At a 3.9:1 to 5.4:1 split, the trainer gets roughly 1,640 to 2,140 of the 10,500 GPUs (reasoned), and the rest generate.

Long rollouts are why staleness is the central problem. A rollout that runs for an hour is generated by several successive checkpoints, so its early tokens come from an older policy than its late ones. Reflection tags every token with the weight version that produced it and trains with asynchronous policy gradients that account for that age. Their figure shows the oldest sample in a batch reaching 107 versions behind, about a day old, while the KL between sampling and updated policy stays bounded. 107 versions in a day is one new version about every 13.5 minutes (reasoned).

Two stacked line charts over training steps 190 to 320. Top: sample age in policy versions behind, with the oldest sample in each batch climbing in a jagged line to a peak of 107 versions, about a day old, while the batch average stays under 10. Bottom: KL divergence between sampling and updated policy, noisy but bounded.
Staleness during RL: the oldest sample in a batch reaches 107 policy versions behind, about a day old, while the batch average stays low and the KL stays bounded. The blog does not describe the algorithm that keeps it stable beyond 'new algorithms' (Reflection, Beam blog, Figure 4).

The algorithm itself is not described. That is the main gap in the RL section. The infrastructure ideas, asynchronous rollout fleets, versioned tokens, fast weight broadcast, are the same family I covered in Fireworks' distributed RL and scaling agentic RL; Reflection's contribution is the scale and the claim that it stayed stable.

Does capability keep improving?

"Across our eval suite, capabilities continued to improve as we increased RL - with no signs of plateau." The evidence is this figure, covering the reasoning expert's training (80M of the 100M+ rollouts).

Three line charts of score against number of rollouts in millions, zero to about 80. DeepSWE rises from about 5% to 41%, HLE from about 20% to 36%, Terminal Bench 2.1 from about 53% to 78%. All three curves are concave and still rising at the last point.
Score against cumulative RL rollouts for Beam's reasoning expert, 80M of the 100M+ rollouts. The blog compares this with 30M rollouts for Inkling and 753K for MiMo (Reflection, Beam blog, Figure 3).

Read it honestly. All three curves are still going up at 80M rollouts, so "no plateau" is true of these curves. They are also clearly concave on a linear axis: DeepSWE gains about 13 points in the first 13M rollouts and about 10 in the last 37M (measured, from the chart). Terminal Bench goes from 70% at about 20M to 78% at 80M. That is the shape of log-linear scaling, the same shape Inkling reported over its 30M rollouts. It is not a plateau; it is a bill that grows exponentially per point.

One more detail is worth noticing. The rollouts plot is for a "reasoning expert", and the safety section explains that the final model was produced by multi-teacher on-policy distillation: one RL teacher for capability, a separate SFT+RL teacher for safety and alignment, merged into Beam. So the 80M-rollout curve is the teacher's, and the released model is a distillation of it. That is a common recipe now; it means the scaling curve and the shipped checkpoint are not the same model.

The compute ledger

Now the cost side. For pretraining, the standard estimate is

C≈6⋅Nactive⋅DC \approx 6 \cdot N_{\text{active}} \cdot D

two FLOPs per weight per token forward, four backward. The thread says "24T" tokens; the blog says 23.8 trillion. With 23B active that is 6×23×109×23.8×1012≈3.3×10246 \times 23\times10^{9} \times 23.8\times10^{12} \approx 3.3\times10^{24} FLOPs (reasoned). Reflection's own pretraining-scaling plot gives a check: I located the "Beam Base" dot on its log axis at about 1024.4710^{24.47}, about 2.9×10242.9\times10^{24} (measured). My 6ND figure is 0.05 decades, about 12%, above it, which is within what one choice of convention (embeddings, attention FLOPs) moves.

Two log-log charts of validation loss in bits per byte against training FLOPs, for code and web data. A dotted line through a series of smaller Reflection runs extends to Beam Base just past 1e24 FLOPs. DeepSeek v4 Flash Base and Nemotron 3 Ultra Base are plotted for comparison.
Pretraining loss against compute, on decontaminated code and web validation sets. Smaller runs predict Beam Base's loss along a straight line over four orders of magnitude; the comparison points are other open base models (Reflection, Beam blog, Figure 6).

The blog also gives the hardware: "under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs", with 92.3% goodput late in the run and nine semi-automatic rewinds. At the full 28 days that is 4.13M GPU-hours, and 3.3×10243.3\times10^{24} FLOPs spread over them is about 221 TFLOP/s sustained per GPU (reasoned). Against the 2,250 TFLOP/s dense BF16 peak that Ai2 normalises B300 MFU to in Olmo-core 3, that is under 10%. Even at three weeks it is about 13%.

That number is low, and I do not think it means the cluster idled. It means 6ND undercounts this run, or the run is not what the four-week figure covers. Attention FLOPs at long context are excluded from 6ND; a fine- grained MoE pays communication costs that never show up as FLOPs; "under four weeks" may be for the final run alone, without midtraining or ablations; and some fraction of 6,144 GPUs may have been spare. The technical report should resolve it. For now, the honest reading is: the token count and active size put Beam at about 3×10243\times10^{24} FLOPs, which is a mid-sized frontier pretraining run, and the GPU count is generous for it.

The RL phase is where the hardware went. 10,500 GPUs for 4 weeks is 7.06M GPU-hours, against at most 4.13M for pretraining (reasoned). RL used about 1.7x the GPU-hours of pretraining. For comparison, the whole Kolibri-1 training run was reported at 392k GPU-hours.

Compute ledger
Kolibri-14.15e23 FLOPs
3.46B active x 20T tokens
Beam3.28e24 FLOPs
23B active x 23.8T tokens
GLM-5.26.84e24 FLOPs
40B active x 28.5T tokens
Inkling1.11e25 FLOPs
41B active x 45T tokens
log axis, 1e23 to about 3e25
implied sustained: 221 TFLOP/s per GPUof a 2,250 TFLOP/s BF16 peak: 9.8%

The ledger puts Beam's 6ND figure beside the other open MoEs on this site that report both active size and token count: Kolibri-1 at about 4.2×10234.2\times10^{23}, GLM-5.2 at 6.8×10246.8\times10^{24}, Inkling at 1.1×10251.1\times10^{25} (all reasoned from their reported numbers). Beam's pretraining is the cheapest of the three frontier-class models by this measure. The GPU-hour view shows the inversion: for Beam, RL is the larger line item.

That inversion is the most interesting thing in the announcement. For years the rule of thumb was that post-training is a rounding error on pretraining. Here a lab reports spending more GPU time on RL than on pretraining, and a capability curve that has not flattened at the end. The RL GPU-hours are not the same as RL FLOPs, since most of those GPUs ran inference at low utilisation waiting on sandboxes, but as a budget line it is the right comparison.

What cannot be checked yet

Everything that matters for using the model:

  1. The architecture numbers. Expert count, top-k, attention window, head layout, vocabulary. Without them no KV-cache sizing, no memory plan, no independent active-parameter count.
  2. The efficiency claim in practice. Tokens per task can be re-measured by anyone with the weights and a harness. Serving cost per solved task, including prefill, cannot until someone runs it.
  3. The benchmarks. Every score here is Reflection's or a third party's that Reflection quoted. The harnesses, effort settings and pass@1 sampling are not documented in the blog.
  4. The quantizations. FP8 and NVFP4 are promised. How much score each gives up is a number that should be in the model card.
  5. The RL algorithm. "New algorithms" for staleness and training-inference mismatch are claimed, not described.

Reflection has said the technical report comes with the weights. If it carries the config and the stability recipe in the detail the blog's charts suggest exists, Beam will be one of the better documented open models. Until then it is a well-made set of charts, and the efficiency chart, at least, says exactly what it measures.

Sources

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Reflection Beam: a 501B open MoE that spends 23B per token, and what its efficiency chart measures", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026reflectionbeam,
  author = {Satyajit Ghana},
  title  = {Reflection Beam: a 501B open MoE that spends 23B per token, and what its efficiency chart measures},
  url    = {https://ai.thesatyajit.com/articles/reflection-beam},
  year   = {2026}
}
share