[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-beam-rl-factory-100m-rollouts":3,"topics-all":38,"news-related-44b118b2-cb3d-47bb-8f26-bbce777cfb31":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"44b118b2-cb3d-47bb-8f26-bbce777cfb31","1 亿 rollout 背后:Beam 的 RL 训练工厂","Reflection AI 首个开放权重模型 Beam:501B 稀疏 MoE、每 token 激活 23B,主打编码与智能体。官方博客罕见公开 RL 基建细节:1.05 万张 GB300 四周生成超 1 亿条 rollout,token 版本化让落后 107 个版本的样本照样稳定训练,权重推送中位 12 秒。","10 月 5 日,隐身很久的 Reflection AI 交出首个开放权重模型 Beam:501B 总参数的稀疏 MoE,每 token 只激活 23B,主打编码、推理与智能体任务,权重本月内以 Apache 2.0 放出,目前还在最后红队评测。跑分之外,这家英伟达支持的初创公司在[官方博客](https:\u002F\u002Freflection.ai\u002Fblog\u002Fintroducing-beam)里写了一份相当少见的「RL 基建说明书」——把强化学习跑到 1 亿条 rollout 的规模,背后是一整套训练工厂。\n\n## 先看规模:1 亿 rollout 意味着什么\n\n官方口径:RL 阶段动用 10,500 张 NVIDIA GB300 连续训练 4 周,生成超过 1 亿条 rollout,单条最长 256K token 上下文,训练与评分共消耗约 13 亿个沙箱,环境池接近 100 万个,全程平均维持 11 万条并发 rollout。Reflection 自称这是开放实验室公开记录里规模最大的 RL 训练之一——注意这是官方说法,目前没有第三方复现。\n\n预训练同样不小:6,144 张 GB300 NVL72 不到 4 周吃完 23.8 万亿 token,官方称训练有效时间占比(goodput)后期达到 92.3%,全程只做了 9 次半自动回滚。\n\n## 异步 RL 的真难题:一天前的数据还能不能学\n\nBeam 用全异步策略梯度训练。规模一大,策略过期(staleness)就是头号不稳定来源:一条长 rollout 里,不同 token 可能出自不同版本的权重,越早的 token 相对当前策略越旧,训练与推理引擎的数值差异还会放大这个问题。Reflection 的做法是给每个 token 标记生成它的权重版本,让算法显式处理过期样本——官方展示的极端情况是,学习落后当前策略 107 个版本(约一天)的数据,数值依然稳定。\n\n配套细节全是工程活:新权重推到整个推理集群,中位只要约 12 秒——先跨机架走 RoCE、机架内走 NVLink 的层级分发,比每个副本各自拉权重少了 75% 的跨机架流量,全网更新快 2.2 倍;训练期间 71 次推理侧故障全部不中断训练任务,中位 8 分钟恢复,损失只占推理 GPU 分钟数的 0.02%;推理与训练的 GPU 配比在 3.9:1 到 5.4:1 之间动态调节,训练器在同一训练谱系里换过 5 种 mesh 配置而不丢状态。沙箱侧峰值 17 万并发,平台累计处理超 10 亿次创建请求,横跨 20 多个集群、两个云、四个区域,90% 的新沙箱 10 秒内就绪;动态 packing 让训练 batch 平均保持 99.99% 填满,平均 rollout 长度涨了近 70%,单卡训练吞吐却只下降 1.5% 以内。\n\n## 成绩单与真实位置\n\n按官方跑分,Beam 在 AIME 2026 拿 97.8,GPQA Diamond 90.5,Terminal Bench 2.1 拿 80.1,SWE Bench Pro v2-Hard 77.2。但放进开源竞技场看更直白:DeepSWE v1.1 上 Beam 44.4,基本追平 GLM 5.2 的 44.0,落后 Qwen 3.8-Max(51.0)、GLM 5.3(61.0)、Kimi K3(68.0)与 DeepSeek V4.1 Flash(74.2)——[中文报道](https:\u002F\u002Fm.toutiao.com\u002Farticle\u002F7693609100252381731\u002F)普遍把这一点读作「仍落后中国头部开源模型」。\n\nBeam 主打的是效率:官方称对标 GLM 5.2 档位的推理成绩,推理算力只需其 1\u002F3 到 1\u002F4;对比 2T+ 参数级的 Qwen 3.8-Max,每 token 推理开销显著更低。它不做最强,做每块钱买到的智能。\n\n## 所以呢\n\n权重开源后人人可下载,但让 RL 稳定吃下 1 亿 rollout 的那套平台——token 版本化、12 秒权重分发、17 万并发沙箱——是组织出来的工程能力,不是 checkpoint 文件。模型会开源,工厂不会。对同样在卷 MoE 和智能体的团队来说,Beam 这份账单的启示比跑分更硬:下一代能力的分水岭,可能就在谁先把 RL 工厂建起来。","https:\u002F\u002Fm.toutiao.com\u002Farticle\u002F7693609100252381731\u002F","b0384237-a532-4411-9665-f40d81178241",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"756111c2-2478-4c71-9a85-4ce64277df1b","en","Inside Beam's RL Factory: 100M Rollouts Without Crashing","Reflection's Beam: 501B sparse MoE, 23B active, for coding and agents. Its RL run: 10,500 GB300 GPUs, 100M+ rollouts, 12-second weight pushes.","On October 5, Reflection AI unveiled Beam, its first open-weight model: a sparse MoE with 501B total parameters and only 23B active per token, built for coding, reasoning and agentic workloads. Weights arrive later this month under Apache 2.0, pending final red-teaming. Beyond the benchmarks, the NVIDIA-backed startup published something rarer in its [official blog](https:\u002F\u002Freflection.ai\u002Fblog\u002Fintroducing-beam): a detailed blueprint of the infrastructure needed to push reinforcement learning past 100 million rollouts.\n\n## The scale: what 100 million rollouts means\n\nPer the official post, the RL stage ran on 10,500 NVIDIA GB300 GPUs for four straight weeks, generating over 100 million rollouts with a 256K-token maximum context. Training and grading consumed roughly 1.3 billion sandboxes, the environment pool held nearly one million tasks, and the run sustained an average of 110K concurrent rollouts. Reflection calls it one of the largest RL runs by any open lab — a vendor claim, not yet independently reproduced.\n\nPretraining was no small feat either: 6,144 GB300 NVL72 GPUs chewed through 23.8 trillion tokens in under four weeks, with the official goodput reaching 92.3% late in the run and only nine semi-automatic rewinds throughout.\n\n## Async RL's real problem: learning from day-old data\n\nBeam was trained with fully asynchronous policy gradients. At this scale, policy staleness becomes the main source of instability: within one long rollout, different tokens may come from different weight versions, and numerical mismatch between training and inference engines compounds the drift. Reflection tags every token with the weight version that produced it, so the algorithm handles stale samples explicitly. In the extreme case the company shows, learning stayed numerically stable even when training on data 107 weight versions (about a day) behind the current policy.\n\nThe supporting details are pure infrastructure work. New weights reach the inference fleet in a median of ~12 seconds, distributed hierarchically over RoCE across racks then NVLink within them — 75% less cross-rack traffic and 2.2× faster fleet-wide adoption than every replica pulling directly. Seventy-one inference incidents during the run were absorbed without killing the training job; capacity recovered in a median of eight minutes, with losses amounting to 0.02% of serving GPU-minutes. The inference-to-training GPU ratio flexed between 3.9:1 and 5.4:1, and the trainer was resized across five mesh configurations within the same lineage without losing state. Sandboxes peaked at 170K concurrent, the platform processed over one billion creation requests across 20+ clusters, two clouds and four regions, with 90% of new sandboxes ready in under 10 seconds. Dynamic packing kept training batches 99.99% full on average while mean rollout length grew almost 70%, holding per-GPU trainer throughput within 1.5%.\n\n## The scorecard and its true position\n\nOn official benchmarks, Beam scores 97.8 on AIME 2026, 90.5 on GPQA Diamond, 80.1 on Terminal Bench 2.1 and 77.2 on SWE Bench Pro v2-Hard. But the open-weight arena tells a blunter story: on DeepSWE v1.1, Beam's 44.4 essentially ties GLM 5.2's 44.0 while trailing Qwen 3.8-Max (51.0), GLM 5.3 (61.0), Kimi K3 (68.0) and DeepSeek V4.1 Flash (74.2). Chinese coverage reads this as still behind the leading Chinese open models.\n\nBeam's pitch is efficiency: the company claims GLM 5.2-tier reasoning performance at one-third to one-quarter of the inference compute, and notably lower per-token cost than 2T+ parameter models like Qwen 3.8-Max. It is not trying to be the strongest — it is selling intelligence per dollar.\n\n## So what\n\nOnce the weights drop, anyone can download the model. But the platform that keeps RL stable across 100 million rollouts — token versioning, 12-second weight distribution, 170K concurrent sandboxes — is organizational engineering, not a checkpoint file. Models get open-sourced; factories do not. For teams also racing on MoE and agents, the real lesson of Beam's bill is harder than its benchmarks: the next capability divide may hinge on who builds the RL factory first.","beam-rl-factory-100m-rollouts","2026-10-07T07:15:00Z","2026-10-07T07:10:49.687259Z","2026-10-07T07:10:49.687276Z",true,"agent",223,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"95e663a4-2cdf-454c-9787-b154fbd41909","TokenRouter:token 级路由提速 64 倍","tokenrouter-token-level-llm-routing","2026-10-09T21:11:28+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"5409a0b8-8d54-4b20-9310-8cde3239b132","小米把推理速度卖成商品:同一权重,十倍价","xiaomi-mimo-ultraspeed-latency-pricing","2026-10-09T15:11:05+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"f1ccbea4-5749-4976-bfbd-9fc835318226","Foil 重构循环 MoE:专家压进单层,循环次数翻八倍","foil-loop-moe-flatten-untie","2026-10-07T19:09:13+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"ccce6dfe-776f-4cd8-9605-9163daea4627","Jeff 开源决策模型:2B 追平 Jev","jeff-open-decision-models-0-8b","2026-09-29T15:10:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"523d3ae4-9c60-4bcb-8e6c-ea98e58c7f71","一次前向一个决策:GLM-5.3-Flash 平替 Jev","glm-flash-single-token-decisions","2026-09-28T13:11:55+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"3559e613-9558-48e1-ab20-f53b62796363","让每个 token 用上全部专家:高德 IntBMoE 解耦参与度、计算与显存,60ms 服务数亿用户","intbmoe-full-participation-block-moe","2026-09-21T13:01:54+00:00"]