[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mask-forcing-video-diffusion-distillation":3,"topics-all":38,"news-related-caed836e-2168-418f-b5c1-bde3ce962e66":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"caed836e-2168-418f-b5c1-bde3ce962e66","Mask Forcing 往蒸馏 rollout 里掺干净 token:修视频生成的模式坍缩,指令遵循最高涨 6.5 分","HKUST 团队提出 Mask Forcing:在 AR 视频扩散蒸馏的 self-rollout 中随机掩码、注入更干净 token,修复 DMD 反向 KL 导致的模式坍缩。三个 baseline 上视觉质量与指令遵循一致提升,frame-wise 动态度最高 +51,推理代码已开源。","实时视频生成的「快」与「好」长期是一对矛盾:蒸馏到少步推理后,画面经常出现过饱和、过平滑,细节像被抹掉一层。9 月 8 日放上 arXiv 的论文 Mask Forcing(arXiv:2609.09123)给出了一个出人意料地简单的修法——不改架构、不动教师模型、不碰推理流程,只在训练时往 rollout 输入里\"掺沙子\"。\n\n## 问题:蒸馏后的学生为什么变\"糊\"\n\n主流路线是把预训练的双向视频扩散模型,通过分布匹配蒸馏(DMD)蒸成因果的自回归(AR)学生,实现实时生成。但团队指出根子在 DMD 的损失函数:反向 KL 具有 mode-seeking(模式追逐)特性,学生分布会坍缩到教师分布的少数几个模式上,宏观表现就是过饱和、过平滑、真实感不足。雪上加霜的是,DMD 目标只在完整的 rollout 输出上评估,不直接约束每个中间转移,误差会沿着去噪步和后续 chunk 逐步累积。\n\n## 方法:双重噪声掩码 rollout\n\nMask Forcing 的核心动作叫 Dual-Noise Masking Rollout:在蒸馏的 self-rollout 过程中,沿空间轴和时间轴做随机掩码,把部分 token 替换成噪声水平更低(更干净)的输入。这一扰动带来两重收益:一是强迫学生 rollout 去探索教师分布的更多区域,让 DMD 的学习信号不再局限于学生已覆盖的模式;二是干净 token 充当去噪引导,帮更噪的 token 做出更好的中间预测、缓解误差累积。论文说明该方法不引入真实视频监督、无额外后训练阶段、无额外网络前向,推理保持不变,可插入 chunk-wise 与 frame-wise 两类现有蒸馏管线。\n\n## 结果:动态度的剧烈分化值得细看\n\n在三个 AR 蒸馏 baseline(Self Forcing、Causal Forcing、LongLive)上,加 Mask Forcing 后视觉质量与指令遵循一致提升:chunk-wise 设定下 Causal Forcing 的 HPSv3 从 9.37 涨到 10.17(+0.80),LongLive 的 Instruct. 分从 42.48 到 42.65;frame-wise 设定下 LongLive 的动态度从 25 跳到 76(+51)。但方向有分化:LongLive 在 chunk-wise 下动态度从 76 掉到 69(−7),frame-wise 下 Causal Forcing 从 28 跳到 52。消融给出的默认操作点:掩码比例 α=0.2、时间步窗口 Δ=250、per-frame + per-chunk 掩码方案。\n\n## 开源与适用性\n\n推理代码与 checkpoint 已在 GitHub 开源(Apache-2.0,基于 Wan2.1-T2V-1.3B 基座,52 星),训练代码标注\"即将放出\"。机构署名 HKUST(GZ)\u002FHKUST、LightSpeed、UCSD、CUHK(SZ)、NUS。\n\n## 所以呢\n\n蒸馏的本质是让一个便宜的学生去逼近昂贵的教师,而反向 KL 的模式追逐意味着学生会\"偷懒\"地只学教师最显眼的几个招式。Mask Forcing 的启示在于:修训练动力学不一定需要更多数据或更大模型,给 rollout 加一点结构化的噪声,让学生的探索轨迹散得更开,反而更接近教师分布的全貌。这个\"加噪修复坍缩\"的思路,和 LLM 侧 RL\u002F蒸馏常见的熵坍缩问题同构——下次看到蒸馏模型输出千篇一律时,值得先查损失函数是不是 mode-seeking 的。\n\n参考:arXiv:2609.09123 · github.com\u002Fdelaprada\u002FMask-Forcing · alicezrzhao.github.io\u002Fmask_forcing\u002F","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.09123","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"604372b6-04e1-4279-9d1e-d12024a3adf0","en","Mask Forcing Fixes Mode Collapse in Video Diffusion Distillation","Injecting cleaner tokens into AR video diffusion rollouts counters mode collapse. Quality rises on three distillation baselines; dynamics jump +51.","Speed and quality have long been at odds in real-time video generation: after distillation to few-step inference, outputs often show over-saturation and over-smoothing. Mask Forcing (arXiv:2609.09123, posted Sep 8) offers a surprisingly simple fix — no architecture change, no teacher modification, no inference-time change, just injecting structured noise into rollout inputs during training.\n\n## The Problem: Why Distilled Students Look \"Washed Out\"\n\nThe mainstream route distills a pretrained bidirectional video diffusion model into a causal autoregressive (AR) student via Distribution Matching Distillation (DMD) for real-time generation. The team traces the root cause to DMD's loss: the reverse KL objective is mode-seeking, so the student distribution collapses onto a few modes of the teacher distribution — manifesting as over-saturation, over-smoothing, and limited realism. Worse, the DMD objective is evaluated only on completed rollout outputs and does not directly regularize each intermediate transition, letting errors accumulate across denoising steps and later chunks.\n\n## The Method: Dual-Noise Masking Rollout\n\nThe core move, Dual-Noise Masking Rollout, applies random masks along spatial and temporal axes during the self-rollout of distillation, replacing some tokens with lower-noise (cleaner) inputs. This yields two benefits: it forces student rollouts to explore more regions of the teacher distribution, so DMD's learning signal is no longer confined to modes the student already covers; and the cleaner tokens act as denoising guidance for noisier ones, improving intermediate predictions and reducing error accumulation. The paper notes the method requires no real-video supervision, no additional post-training stages, no extra network forward passes, and leaves inference unchanged — pluggable into both chunk-wise and frame-wise distillation pipelines.\n\n## Results: The Sharp Divergence in Dynamics Deserves Scrutiny\n\nAcross three AR distillation baselines (Self Forcing, Causal Forcing, LongLive), adding Mask Forcing consistently improves visual quality and instruction following: in the chunk-wise setting, Causal Forcing's HPSv3 rises from 9.37 to 10.17 (+0.80), and LongLive's Instruct. score from 42.48 to 42.65; in the frame-wise setting, LongLive's Dynamic Degree jumps from 25 to 76 (+51). But note the divergence: under chunk-wise, LongLive's dynamics drop from 76 to 69 (−7), while under frame-wise, Causal Forcing's dynamics jump from 28 to 52. The ablations' default operating point: mask ratio α=0.2, timestep window Δ=250, per-frame + per-chunk masking.\n\n## Open Source and Applicability\n\nInference code and checkpoints are open-sourced on GitHub (Apache-2.0, built on the Wan2.1-T2V-1.3B base, 52 stars), with training code marked \"coming soon\". Institutional affiliations: HKUST(GZ)\u002FHKUST, LightSpeed, UCSD, CUHK(SZ), and NUS.\n\n## So What\n\nDistillation is fundamentally about a cheap student approximating an expensive teacher, and reverse KL's mode-seeking means the student \"cheats\" by learning only the teacher's most conspicuous tricks. The lesson of Mask Forcing: fixing training dynamics doesn't necessarily require more data or bigger models — adding structured noise to rollouts to spread student exploration trajectories can get you closer to the full teacher distribution. This \"fixing collapse with noise\" idea is structurally similar to entropy collapse in LLM-side RL\u002Fdistillation — the next time a distilled model produces uniform outputs, check whether the loss is mode-seeking first.\n\nRefs: arXiv:2609.09123 · github.com\u002Fdelaprada\u002FMask-Forcing · alicezrzhao.github.io\u002Fmask_forcing\u002F","mask-forcing-video-diffusion-distillation","2026-09-09T23:08:37Z","2026-09-09T23:08:40.668042Z","2026-09-09T23:08:40.668055Z",true,"agent",112,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"02c8b500-ec11-44a6-8c58-6e880563dad8","FastH3 开源:4 步蒸馏版 MiniMax H3,B200 单卡最高提速 14 倍","fasth3-4-step-distilled-minimax-h3","2026-08-30T21:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"18d2aa73-7244-4b10-b611-46475e17327e","ForgeWM开源:一步去噪72FPS的可玩世界模型,8张卡复现全流程","forgewm-few-step-playable-world-model","2026-08-24T21:10:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"0599b775-ac17-49d2-aebd-a16f531c7168","腾讯混元 MeanFlowNFT：把 RL 接进「平均速度生成器」，Wan 2.1 4 步反超 50 步 LongCat-Video RL","tencent-hunyuan-meanflownft","2026-07-16T12:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"42cfc778-8f1b-4bf2-a0ae-4343a066f48d","RhymeFlow：清华提出异步去噪流调度，DiT视频生成训练免费加速1.53倍","rhymeflow-tsinghua-async-denoising-1-53x","2026-06-07T22:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"bf8755fc-cd4f-4bd9-9617-e70f56ddc4ac","LTX-2.3：开源视频生成正式进入 4K + 原生音频时代","ltx-2-3-lightricks-4k-native-audio","2026-06-02T01:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"fb2da954-b12e-4bde-9146-b61dd240df92","SolarWM 开源:143 万条视频喂出的世界模型,5 秒训练片段撑起小时级交互","solarwm-open-data-video-world-models","2026-09-03T15:08:13+00:00"]