[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ultradata-rl-2609-verifiable-rl-dataset":3,"topics-all":38,"news-related-e84fe968-5d86-4247-baad-5da23efef860":39},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"e84fe968-5d86-4247-baad-5da23efef860","UltraData-RL-2609 开源:85,995 条可验证奖励任务,拆解 MiniCPM5-2B 的 RL 燃料","OpenBMB 开源 MiniCPM5-2B 后训练 RL 数据集 UltraData-RL-2609:85,995 条任务覆盖数学、代码、长上下文与 STEM 四域,每条带机器可判定的 ground truth。数据集卡记录 AIME 2025 在约 300 个 RL step 内从 61 升至 81。","小模型圈这两周的主线是「2B 打 4B」,OpenBMB 现在把另一半故事也放了出来:MiniCPM5-2B 后训练烧掉的 RL 燃料——UltraData-RL-2609 数据集已完整开源。模型权重 9 月 7 日落地,数据集同日跟上,「权重 + 数据」整套放出,这种做法在端侧模型厂商里并不多见。\n\n这份语料是 UltraData L0-L4 分层数据框架中的 L3 精炼层,专门服务 RL 阶段:85,995 条任务,全部是可验证奖励(verifiable reward)型——每条都带机器能判对错的 ground truth,不依赖 LLM 裁判的自由心证。\n\n## 四个方向,每条都要可判\n\n分布相当均衡:数学 32,412 条(37.7%)、代码 23,665 条(27.5%)、长上下文 18,046 条(21.0%)、STEM 知识 11,872 条(13.8%)。验证方式按域分化:数学与知识题比对标准答案;代码任务在测试用例上实际执行,stdin\u002Fstdout 逐例比对;长上下文则保证上下文一定支持答案。选择题、判断题、证明题、多小题与依赖图片的条目,在构建管线里被直接剔除——机器判不了的东西不进 RL。\n\n## 六道工序,专治标签不可信\n\n管线重心在第 4、5 阶段。第 4 阶段查奖励可靠性:LLM 裁判审查问答一致性,数学\u002F知识答案由多个独立模型重解、按共识定标签,凑不出共识的直接丢弃,不猜。第 5 阶段做难度校准:在 RL 初始 checkpoint 上反复 rollout 估计每题通过率,已完全掌握的(通过率 1)删掉,可学习区间保留,通过率为 0 但标签确认有效的硬题也保留,交给在线动态采样;该阶段只改难度与采样权重,永不改参考标签。上游来源全部是公开数据集的加工重组:数学取 DAPO、DeepScaleR、DeepMath 的并集,知识取 OpenScienceReasoning-2,长上下文基于 HotpotQA、Qasper、MuSiQue 加长并补充自造数据,代码取 OpenCodeReasoning 系列加 HardTests。\n\n## 效果实录与三处边界\n\n数据集卡记录:JustRL II 实验设置下,AIME 2025 在约 300 个 RL step 内从 61 升到 81;消费该语料完成 post-training 的最终 checkpoint 在 AIME 2025 拿到 86.5(模型卡评测表),表内平均分 53.9,高于对比集中列出的 4B 级模型(最高 51.1)。配套 RL+OPD 阶段平均拉升推理与通用能力 10.96 分、agentic 能力 6.96 分。\n\n边界同样写得明白:代码域只给测试用例不给沙箱,奖励要使用者自己算;难度过滤与在线采样权重没有存为发布字段;去污染只覆盖构建时已知的评测集,引入新 benchmark 前要重查。许可证还有个反直觉条款:Apache 2.0 之下明确禁止未经书面许可的原样转存、镜像与商业再打包——「开源」与「随意搬运」被切开对待。\n\n对做小模型 RL 的团队,这份 188 GB 的语料是少有的可整条复用的工程参考:端侧 2B 的逆袭不只靠架构,更靠「每条奖励都可信」的数据纪律。下一个问题是,同样的管线搬到 agentic 数据上,还能不能维持这个验证密度。\n\n参考:[UltraData-RL-2609 数据集卡](https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002Fopenbmb\u002FUltraData-RL-2609) \u002F [MiniCPM5-2B 模型卡](https:\u002F\u002Fhuggingface.co\u002Fopenbmb\u002FMiniCPM5-2B)","https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002Fopenbmb\u002FUltraData-RL-2609","71df6775-935e-4b09-bda9-e03ee3eb8191",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"22c44a1a-60ad-4d6f-be17-d3e5c7259275","en","OpenBMB opens UltraData-RL-2609: 85,995 verifiable RL tasks","OpenBMB released the RL corpus behind MiniCPM5-2B: 85,995 verifiable tasks across math, code, long-context and STEM, built via a six-stage pipeline.","The small-model storyline this month is \"2B beats 4B,\" and OpenBMB has now published the other half: UltraData-RL-2609, the RL corpus burned during MiniCPM5-2B's post-training, is fully open. The model weights landed on September 7 and the dataset arrived the same day — weights plus data in one drop, still rare among on-device vendors.\n\nThe corpus is the L3 refined layer of the UltraData L0-L4 tiered data framework, built for the RL stage: 85,995 tasks, all verifiable-reward — each carries a ground truth a machine can grade, with no free-form LLM judging.\n\n## Four directions, every item machine-checkable\n\nThe mix is balanced: 32,412 math items (37.7%), 23,665 code (27.5%), 18,046 long-context (21.0%) and 11,872 STEM knowledge (13.8%). Verification differs by domain: math and knowledge answers are matched against reference answers; code submissions are executed against test cases with stdin\u002Fstdout compared case by case; long-context items guarantee the context supports the answer. Multiple-choice, true\u002Ffalse, proof, multi-part and image-dependent items were removed from the pipeline outright — what a machine cannot grade does not enter RL.\n\n## Six stages, aimed at untrustworthy labels\n\nStages 4 and 5 carry the weight. Stage 4 audits reward reliability: an LLM judge reviews question-answer consistency, and math\u002Fknowledge answers are re-solved by several independent models and relabeled by consensus — items without consensus are dropped, not guessed. Stage 5 calibrates difficulty: repeated rollouts on the RL initialization checkpoint estimate a per-item pass rate; fully mastered items (pass rate 1) are removed, the learnable band is kept, and hard-but-valid items (pass rate 0 with a confirmed label) are retained for online dynamic sampling. This stage changes difficulty and sampling weights only — never the reference label. Upstream sources are all public datasets, reworked: math unions DAPO, DeepScaleR and DeepMath; knowledge draws from OpenScienceReasoning-2; long-context extends HotpotQA, Qasper and MuSiQue with longer contexts plus in-house synthetic data; code builds on the OpenCodeReasoning series plus HardTests.\n\n## Measured gains, and three boundaries\n\nThe dataset card records that under the JustRL II setting, AIME 2025 rose from 61 to 81 within about 300 RL steps; the final checkpoint that consumed this corpus in post-training reaches 86.5 on AIME 2025 (model card eval table), averaging 53.9 across that table — above the 4B-class models listed for reference (their best is 51.1). The accompanying RL+OPD stage adds an average of 10.96 points on reasoning and general capabilities and 6.96 on agentic tasks.\n\nThe boundaries are stated plainly: the code domain ships test cases but no sandbox, so users compute rewards themselves; difficulty filters and online sampling weights are not stored as release fields; decontamination covered only benchmarks known at construction time, and a new check is required before introducing a new benchmark. The license adds a counterintuitive clause: under Apache 2.0, unauthorized unchanged re-hosting, mirroring and commercial repackaging are explicitly prohibited — \"open\" and \"repost freely\" are treated as different things.\n\nFor teams doing small-model RL, this 188 GB corpus is a rare end-to-end engineering reference: the 2B comeback rests not only on architecture but on the discipline of making every reward trustworthy. The next question is whether the same pipeline can hold this verification density on agentic data.\n\nRefs: [UltraData-RL-2609 dataset card](https:\u002F\u002Fhuggingface.co\u002Fdatasets\u002Fopenbmb\u002FUltraData-RL-2609) \u002F [MiniCPM5-2B model card](https:\u002F\u002Fhuggingface.co\u002Fopenbmb\u002FMiniCPM5-2B)","ultradata-rl-2609-verifiable-rl-dataset","2026-09-07T23:07:45Z","2026-09-07T23:07:47.761440Z","2026-09-07T23:07:47.761452Z",true,"agent",30,[],{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"2b37a19b-1dde-4238-bef5-39b1d19157f1","OpenBMB 开源 MiniCPM5-2B:2B 端侧模型平均分超对比集 4B 级","openbmb-minicpm5-2b-on-device","2026-09-07T17:02:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"0c29e1ad-914a-4b79-a153-445c087acb03","被 LLM 抛弃的 dropout 翻身:Cerebras 称调好可省 25% 训练 FLOPs","dont-drop-dropout-layer-sparsity","2026-09-07T21:06:35+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"005557c5-8a3c-4d34-89bc-35d5351c4570","蒸馏只需要一条训练样本?清华实测:单条query覆盖71.5%训练状态,16条追平17k全量","one-shot-opd-single-query-distillation","2026-09-05T21:07:11+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"7623f190-7071-4811-a6f1-32462a99b8d3","经验会过期:阿里云论文让自主后训练的有害授权率从 62.5% 降到 25%","bcit-conditional-experience-transfer-post-training","2026-09-05T17:11:11+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"199cd4ef-f092-45a5-8635-91778dd2bce2","编译即训练：一句规约炼出 83.6% 准确率的本地神经函数，教师模型只用一次","compile-by-training-neural-functions","2026-09-04T23:08:03+00:00"]