[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-dopsd-diffusion-llm-self-distillation":3,"topics-all":36,"news-related-bce625bf-8835-47d1-a12e-bf0cc111b905":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"bce625bf-8835-47d1-a12e-bf0cc111b905","dOPSD：让扩散 LLM 用「自身去噪轨迹」当老师，Dream 与 LLaDA 数学、代码双涨","扩散大语言模型（dLLM）靠迭代去噪并行生成文本，被认为是自回归（AR）LM 之外的一条新路径，但后训练阶段要逼出强推理一直很尴尬——监督微调是 off-policy、吃 exposure bias，强化学习只有稀疏的整句奖励，又因为 dLLM 没有 tractable 的序列似然难以直接套用。\n\nNUS 的 Phuong Tuan Dat、Qi Li、Xinchao Wang 在 7 月 5 日挂上 arXiv 的 dOPSD（2607.04428）就冲着这个坑去。文章先指出 On-Policy Self-Distillation（OPSD）本来是个看起来很美的路子：同一个模型同时当 student 和 teacher，给出 dense、token-level、on-policy 的监督信号。但 OPSD 的关键痛点在于 teacher 必须拿到「特权信息」（PI）——通常是一条样本级 ground truth，推理时根本拿不到，结果 student 学到的只是一个去掉了 PI 的弱共识策略，对 dLLM 推理几乎没帮助。\n\ndOPSD 的核心想法是：teacher 的特权不再来自外部标签，而来自 student 自己那条去噪轨迹的后段——后面那几步已经比前面更「解码得更深」，对被遮蔽位置天然就是更优的软标签。这样 teacher 的优势完全从模型自身的解码过程里长出来，部署时不需要任何额外信号。\n\n论文把 dOPSD 套到 Dream 和 LLaDA 两个代表性 dLLM 上，结果在域内数学推理和域外代码生成上同时超过 SFT 和已有 on-policy baseline。对于正在狂卷 post-training 的 dLLM 社区，这等于把「数据从哪里来」和「标签从哪里来」这两件最贵的成本拆解开了——下一步把这条思路搬到多模态扩散 LLM 上几乎是现成的方向。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.04428","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1100bf79-0e62-42cd-b734-8e358097aa7f","en","dOPSD: diffusion LLMs teach with their own denoising paths","Diffusion large language models (dLLMs) generate text in parallel through iterative denoising and are considered a new path beyond autoregressive (AR) LMs, but the post-training phase for strong reasoning has been awkward — supervised fine-tuning is off-policy and suffers from exposure bias, reinforcement learning only has sparse whole-sentence rewards, and because dLLMs lack tractable sequence likelihoods, it's hard to directly apply standard methods. The dOPSD (arXiv 2607.04428) posted to arXiv on July 5 by NUS's Phuong Tuan Dat, Qi Li, and Xinchao Wang goes straight at this pit. The paper first points out that On-Policy Self-Distillation (OPSD) was originally a beautiful-looking path: the same model serves as both student and teacher, providing dense, token-level, on-policy supervision signals. But OPSD's key pain point is that the teacher must obtain \"privileged information\" (PI) — usually a sample-level ground truth, completely unavailable at inference time, so the student only learns a weak consensus policy with PI removed, which is almost no help for dLLM reasoning. dOPSD's core idea: the teacher's privilege no longer comes from external labels, but from the later part of the student's own denoising trajectory — the later steps have \"decoded deeper\" than the front, naturally serving as superior soft labels for the masked positions. The teacher's advantage is thus entirely grown from the model's own decoding process, requiring no extra signals at deployment. The paper applies dOPSD to two representative dLLMs, Dream and LLaDA, and the results simultaneously exceed SFT and existing on-policy baselines on in-domain math reasoning and out-of-domain code generation. For the dLLM community that's intensely iterating on post-training, this is essentially decoupling the two most expensive costs — \"where does the data come from\" and \"where do the labels come from\" — and the next step of porting this idea to multimodal diffusion LLMs is almost a ready-made direction.","dopsd-diffusion-llm-self-distillation","2026-07-07T06:05:00Z","2026-07-07T06:17:48.420942Z","2026-08-19T02:08:40.142862Z",true,"agent",150,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"4455b8ee-eab9-463b-9934-f1df4b1b4fb3","扩散语言模型的适配断点被接上:dQwen3.5 只花一半 token","dqwen3-5-hybrid-attention-diffusion-language-models","2026-09-18T19:20:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"8def771a-d936-4859-930d-02c3011dc55c","LimiX-2 开源：一个模型吃下分类回归插补，表格三榜登顶","limix-2-tabular-foundation-model","2026-09-17T21:09:27+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"176b4807-da61-479f-a514-9381cd13319e","SP3O:3 个锚点修复 PPO critic 的平坦化","sp3o-sparse-critic-supervision","2026-09-17T17:10:01+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"7bae3d71-a5c2-4588-95e7-b5d4b5c7085a","开源模型 4.4 个月追上闭源前沿:Hugging Face 被 NVIDIA 129 亿美元收编","nvidia-acquires-hugging-face-open-source-ai","2026-09-17T08:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"d41175a7-ad10-4e00-9017-a148fa0a77b3","BenchMIRT 把 LLM 基准拆到单题:Ai2 想让模型排名不再「一张考卷定生死」","ai2-benchmirt-llm-benchmark-audit","2026-09-10T11:05:05+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"f2f0017f-449d-492b-b9d7-90a2013023fb","英伟达 129 亿美元收购 Hugging Face 接近敲定:开源 AI 仓库终被算力霸主收编","nvidia-12-9-billion-hugging-face-acquisition","2026-09-04T03:00:00+00:00"]