On October 5, Reflection AI unveiled Beam, its first open-weight model: a sparse MoE with 501B total parameters and only 23B active per token, built for coding, reasoning and agentic workloads. Weights arrive later this month under Apache 2.0, pending final red-teaming. Beyond the benchmarks, the NVIDIA-backed startup published something rarer in its official blog: a detailed blueprint of the infrastructure needed to push reinforcement learning past 100 million rollouts.
The scale: what 100 million rollouts means
Per the official post, the RL stage ran on 10,500 NVIDIA GB300 GPUs for four straight weeks, generating over 100 million rollouts with a 256K-token maximum context. Training and grading consumed roughly 1.3 billion sandboxes, the environment pool held nearly one million tasks, and the run sustained an average of 110K concurrent rollouts. Reflection calls it one of the largest RL runs by any open lab — a vendor claim, not yet independently reproduced.
Pretraining was no small feat either: 6,144 GB300 NVL72 GPUs chewed through 23.8 trillion tokens in under four weeks, with the official goodput reaching 92.3% late in the run and only nine semi-automatic rewinds throughout.
Async RL's real problem: learning from day-old data
Beam was trained with fully asynchronous policy gradients. At this scale, policy staleness becomes the main source of instability: within one long rollout, different tokens may come from different weight versions, and numerical mismatch between training and inference engines compounds the drift. Reflection tags every token with the weight version that produced it, so the algorithm handles stale samples explicitly. In the extreme case the company shows, learning stayed numerically stable even when training on data 107 weight versions (about a day) behind the current policy.
The supporting details are pure infrastructure work. New weights reach the inference fleet in a median of ~12 seconds, distributed hierarchically over RoCE across racks then NVLink within them — 75% less cross-rack traffic and 2.2× faster fleet-wide adoption than every replica pulling directly. Seventy-one inference incidents during the run were absorbed without killing the training job; capacity recovered in a median of eight minutes, with losses amounting to 0.02% of serving GPU-minutes. The inference-to-training GPU ratio flexed between 3.9:1 and 5.4:1, and the trainer was resized across five mesh configurations within the same lineage without losing state. Sandboxes peaked at 170K concurrent, the platform processed over one billion creation requests across 20+ clusters, two clouds and four regions, with 90% of new sandboxes ready in under 10 seconds. Dynamic packing kept training batches 99.99% full on average while mean rollout length grew almost 70%, holding per-GPU trainer throughput within 1.5%.
The scorecard and its true position
On official benchmarks, Beam scores 97.8 on AIME 2026, 90.5 on GPQA Diamond, 80.1 on Terminal Bench 2.1 and 77.2 on SWE Bench Pro v2-Hard. But the open-weight arena tells a blunter story: on DeepSWE v1.1, Beam's 44.4 essentially ties GLM 5.2's 44.0 while trailing Qwen 3.8-Max (51.0), GLM 5.3 (61.0), Kimi K3 (68.0) and DeepSeek V4.1 Flash (74.2). Chinese coverage reads this as still behind the leading Chinese open models.
Beam's pitch is efficiency: the company claims GLM 5.2-tier reasoning performance at one-third to one-quarter of the inference compute, and notably lower per-token cost than 2T+ parameter models like Qwen 3.8-Max. It is not trying to be the strongest — it is selling intelligence per dollar.
So what
Once the weights drop, anyone can download the model. But the platform that keeps RL stable across 100 million rollouts — token versioning, 12-second weight distribution, 170K concurrent sandboxes — is organizational engineering, not a checkpoint file. Models get open-sourced; factories do not. For teams also racing on MoE and agents, the real lesson of Beam's bill is harder than its benchmarks: the next capability divide may hinge on who builds the RL factory first.