[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openbmb-minicpm5-2b-on-device":3,"news-related-2b37a19b-1dde-4238-bef5-39b1d19157f1":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"2b37a19b-1dde-4238-bef5-39b1d19157f1","OpenBMB 开源 MiniCPM5-2B:2B 端侧模型平均分超对比集 4B 级","OpenBMB 发布端侧模型 MiniCPM5-2B:25.2 亿参数 dense Transformer,128K 上下文,官方 34 项评测平均 53.9 超过对比集内 4B 级,SWE-bench Verified 达 46.4;RL 教师+在线蒸馏配方与训练数据同步开源,Apache-2.0。","端侧大模型又添一员。OpenBMB 联合清华大学 NLP 实验室与 ModelBest,发布了 MiniCPM5 系列的第二款模型 MiniCPM5-2B:一个 25.2 亿参数的 dense Transformer(总参数 2,516,756,480,非嵌入参数约 19.8 亿),42 层,GQA 注意力(16 个查询头、2 个 KV 头),原生 128K 上下文,面向手机、笔记本等资源受限的本地部署,Apache-2.0 协议开源。\n\n## 2B 打 4B:数字怎么读\n\n官方在 34 项基准上给出平均分 53.9,对比集里 4B 级最高是 Qwen3.5-4B 的 51.1——2B 的平均分压过所有更大的对比模型。分项看更明显:AIME 2025 拿到 86.5(Qwen3.5-4B 为 78.8),LiveCodeBench v6 拿到 69.1(同对比 56.4),长上下文 NoLiMa 68.1,中文搜索智能体 BrowseComp-ZH 43.5。最扎眼的是 SWE-bench Verified 的 46.4——真实仓库级编程修复超过对比集内全部 4B 级模型(LFM2.5-2.6B 仅 6.0)。需要说明,这些是官方评测集内的自报数字,部分分数带 Artificial Analysis 官方发布标记。\n\n## RL 教师 + 在线策略蒸馏\n\n训练分三段:base 训练、mid-training、post-training。post-training 先用 400B tokens 的深度思考 SFT 打底,再为数学、代码、agentic、写作等领域训练专门的 RL 教师,最后用 On-Policy Distillation(OPD)把 16 个 RL 专家(含 5 个 agentic 专家)蒸馏进一个发布模型。官方称 RL+OPD 阶段为推理与通用能力平均带来 10.96 分提升,agentic 能力 6.96 分。数据侧同步开源:UltraX 预训练网页数据、UltraData-Code 分层代码数据、50 万条 agent SFT 样本与 8 万余条 RL 样本全部放出。\n\n## 部署生态铺得很开\n\n模型用标准 LlamaForCausalLM 架构,主流推理引擎直接加载,无需自定义 kernel 或模型代码 fork。权重提供 GGUF、MLX、GPTQ、DSpark 草稿模型等多种格式,llama.cpp、Ollama、LM Studio、vLLM、SGLang、ArcLight 全部有官方 cookbook。SGLang 在发布当天给出 day-0 支持,并报告在 RTX 5090 上开启 DSpark 投机解码后,单用户解码速度超过每秒 250 token。国内芯片侧,FlagOS 平台把模型适配到了英伟达、昇腾、摩尔线程、昆仑芯等 9 种芯片。\n\n## 所以呢\n\n两个背景值得注意。一是 Artificial Analysis 刚在 9 月 4 日发布 Intelligence Index v4.2,加大了私有测试集权重,MiniCPM5-2B 随即登顶该指数 4B 以下开源模型;二是 cryptobriefing 提醒,1B 版本在旧指数拿到的 17.9 与 2B 版本的新指数成绩不可直接比较——方法论变了,跨版本分数没有可比性。对开发者,真正有信号的是 SWE-bench 这类真实任务分数:2B 模型开始够到一年前 4B 级都碰不到的线,端侧跑 agent 的实用窗口正在打开。但 benchmark 领先与真实效用并不总同步,值不值得换,得在自己的设备上跑一遍。\n\n参考:官方模型卡(huggingface.co\u002Fopenbmb\u002FMiniCPM5-2B),cryptobriefing 与 SGLang 的报道。","https:\u002F\u002Fhuggingface.co\u002Fopenbmb\u002FMiniCPM5-2B","71df6775-935e-4b09-bda9-e03ee3eb8191",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"eda32d0d-5849-4215-9da5-a1178d05536e","en","OpenBMB's MiniCPM5-2B: 2B on-device model beats 4B-class scores","OpenBMB's 2B on-device model averages 53.9 over 34 benchmarks and beats compared 4B models; weights, data and recipe are open.","On-device LLMs gain another contender. OpenBMB, together with the Tsinghua University NLP lab and ModelBest, has released MiniCPM5-2B, the second model in the MiniCPM5 series: a dense 2B Transformer with 2,516,756,480 total parameters (about 1.98B non-embedding), 42 layers, GQA attention with 16 query heads and 2 KV heads, and a native 131,072-token context window, built for phones, laptops and other resource-constrained local deployment. Weights are open under Apache-2.0.\n\n## 2B versus 4B: how to read the numbers\n\nAcross 34 benchmarks the official average is 53.9, while the best 4B-class model in the comparison set, Qwen3.5-4B, sits at 51.1 — the 2B model outscored every larger model in the set. The per-task gaps are wider: 86.5 on AIME 2025 (Qwen3.5-4B: 78.8), 69.1 on LiveCodeBench v6 (56.4), 68.1 on NoLiMa for long context, and 43.5 on BrowseComp-ZH. The standout is 46.4 on SWE-bench Verified — above every 4B-class model compared (LFM2.5-2.6B managed 6.0). These are vendor-reported numbers within the official comparison set, and some scores carry the official Artificial Analysis release mark.\n\n## RL teachers plus on-policy distillation\n\nTraining runs in three stages: base training, mid-training, then post-training. Post-training starts with 400B tokens of deep-thinking SFT, then trains specialized RL teachers for math, code, agentic tasks and writing, and finally uses On-Policy Distillation (OPD) to distill 16 RL experts, including 5 agentic experts, into one release model. OpenBMB reports the RL+OPD stage added an average 10.96 points to reasoning and general capability, and 6.96 to agentic capability. The data is open too: UltraX web pre-training data, tiered UltraData-Code, 500K agent SFT samples and 80K+ RL samples are all released.\n\n## A wide deployment footprint\n\nThe model uses the standard LlamaForCausalLM architecture, so mainstream inference engines load it directly — no custom kernels, no model-code fork. Weights ship in GGUF, MLX, GPTQ and DSpark draft-model formats, with official cookbooks for llama.cpp, Ollama, LM Studio, vLLM, SGLang and ArcLight. SGLang provided day-0 support and reported over 250 tokens per second per user on an RTX 5090 with DSpark speculative decoding. On the chip side, the FlagOS platform adapted the model to 9 AI chips including Nvidia, Ascend, Moore Threads and Kunlunxin.\n\n## So what\n\nTwo pieces of context matter. First, Artificial Analysis released Intelligence Index v4.2 on September 4, three days before the model, with heavier weight on private test sets; MiniCPM5-2B now leads that index among open models under 4B parameters. Second, cryptobriefing notes the 1B model scored 17.9 on the old index while the 2B version lands on the new one — the methodology changed, so cross-version scores are not comparable. For developers, the real signal sits in task-level scores like SWE-bench: a 2B model reaching territory that 4B-class models could not a year ago opens a practical window for on-device agents. But benchmark leads and real-world utility do not always move together — whether it is worth switching is best answered by running it on your own device.\n\nReference: the official model card (huggingface.co\u002Fopenbmb\u002FMiniCPM5-2B), plus coverage from cryptobriefing and SGLang.","openbmb-minicpm5-2b-on-device","2026-09-07T17:02:00Z","2026-09-07T17:06:34.425682Z","2026-09-07T17:06:34.425690Z",true,"agent",29,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"e84fe968-5d86-4247-baad-5da23efef860","UltraData-RL-2609 开源:85,995 条可验证奖励任务,拆解 MiniCPM5-2B 的 RL 燃料","ultradata-rl-2609-verifiable-rl-dataset","2026-09-07T23:07:45+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"453ce9a1-5d55-4981-b44d-c261b8051724","GLM-5.3 753B 权重上架 HuggingFace,智谱兑现两周开源承诺","glm-5-3-weights-huggingface-release","2026-08-28T15:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00"]