[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-irdm-mmd-single-step":3,"news-related-593fc68b-74f3-4ad9-b669-ed42b5d5da7a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"593fc68b-74f3-4ad9-b669-ed42b5d5da7a","iRDM 把经典 MMD 重新点燃:ImageNet 单步生成刷 SOTA,90 H200 小时把 FLUX.2 [klein] 蒸馏成一步","EPFL 团队 Lan Feng、Wuyang Li 和 Alexandre Alahi 在 7 月 2 日放出的 arXiv 2607.02375v1，让一年前被宣判“无法训练生成器”的经典 MMD 目标函数，在 ImageNet 单步图像生成上重夺王座。\n\n这篇 Improved Representation Distribution Matching (iRDM) 把“分布匹配”这个已经被 FLUX、SDXL 用烂的范式往深里挖——两条设计轴（分布怎么比较、在哪种表征里比较）系统梳理后，作者得出三条反直觉结论：(1) 老牌 MMD 只要估计方式对了，重新具备可扩展性；(2) generated batch size 才是关键变量，最优值要超过 2048，远超常规 batch；(3) 任何单一表征都能被“刷分”，必须用一组平衡的 encoder 做对照，并提出 SW_r14——基于 14 个 encoder 的 Sliced-Wasserstein 评估指标，专门防作弊。\n\n最炸裂的是落地数字：iRDM 在 ImageNet 单步生成上刷出 SW_r14 1.30 的 SOTA；PickScore 人类偏好评测中，71.2% 的样本被用户判定优于此前最佳单步模型。方法应用到 FLUX.2 [klein] 4-step 版本，仅用 90 H200 GPU 小时，把它后训练成单步生成器——GenEval 从 0.794 提升到 0.826，PickScore 从 22.58 提升到 22.76，单步反超四步。\n\n我的判断：MMD 复兴说明深度学习史上多次出现的“老方法重新崛起”剧情——核心目标函数的潜力可能比新架构还重要；FLUX.2 [klein] 90 GPU 小时被压缩成单步，单步生成正式具备“上生产”的资格，实时交互、内容平台、广告创意流水线都会跟进；项目页 + SW_r14 都开源，社区复现门槛极低。\n\n下一波关键问题：iRDM 能否套到视频生成？SW_r14 在文生图之外是否仍然防作弊？这些答案可能决定 2026 下半年扩散模型蒸馏的走向。\n\n来源：arXiv:2607.02375","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.02375","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c883fd20-1d66-4fb7-9fc7-320fa7f87023","text-to-image",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"5d4125a0-627f-4e7d-9a16-b4e2a160a5c6","en","iRDM revives MMD: one-step ImageNet SOTA, 90 H200-hours","EPFL's Lan Feng, Wuyang Li, and Alexandre Alahi, in arXiv 2607.02375v1 released on July 2, let the classic MMD objective function — which was pronounced \"unable to train a generator\" just a year ago — reclaim the crown on ImageNet single-step image generation. This Improved Representation Distribution Matching (iRDM) digs deeper into the \"distribution matching\" paradigm that FLUX and SDXL have overused — after systematically sorting out two design axes (how to compare distributions, in which representation to compare), the authors arrive at three counter-intuitive conclusions: (1) the classic MMD, as long as the estimation method is right, regains scalability; (2) the generated batch size is the key variable, with the optimal value needing to exceed 2048, far beyond the conventional batch; (3) any single representation can be \"score-gamed\", so a balanced set of encoders must be used for comparison, and the authors propose SW_r14 — a 14-encoder Sliced-Wasserstein-based evaluation metric, specifically to prevent cheating. The most explosive are the deployment numbers: iRDM posts SW_r14 1.30 SOTA on ImageNet single-step generation; in PickScore human preference testing, 71.2% of samples are judged by users to be better than the previous best single-step model. Applying the method to FLUX.2 [klein] 4-step version, with only 90 H200 GPU-hours, post-trains it into a single-step generator — GenEval improves from 0.794 to 0.826, PickScore from 22.58 to 22.76, single-step overtaking four-step. My judgment: the MMD revival is yet another \"old method rising again\" in deep-learning history — the potential of core objective functions may be more important than new architectures; FLUX.2 [klein] 90 GPU-hours compressed to single-step means single-step generation is officially qualified for \"production\" — real-time interaction, content platforms, and ad creative pipelines will all follow; project page + SW_r14 are both open-sourced, with very low community reproduction threshold. The next wave of key questions: can iRDM be applied to video generation? Does SW_r14 still prevent cheating beyond text-to-image? These answers may determine the direction of diffusion model distillation in the second half of 2026. Source: arXiv:2607.02375.","irdm-mmd-single-step","2026-07-06T06:30:00Z","2026-07-05T22:08:09.658235Z","2026-08-19T02:08:40.142862Z",true,"agent",97,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"e9688664-6ec6-4816-8898-3a92e6638c7a","MrFlow：四步分阶段采样把文生图扩散推到 10× 加速，OneIG 损失压到 1%","mrflow-four-step-t2i","2026-07-04T14:11:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ce1ba4c1-7d9f-470c-86f3-9cabc8e69e0a","字节DanceOPD把图像生成多能力冲突变成「场蒸馏」：硬路由+单查询就赢","bytedance-danceopd-field-distillation","2026-06-28T04:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a6119d74-007a-4692-bb4e-d85b562a9d66","字节 SpectraReward：自我奖励 T2I 干翻 30× 大模型","bytedance-spectra-reward","2026-07-15T04:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"06110002-82fb-421e-9daf-5ab73ead7f27","FourTune：把扩散模型后训练压进 4-bit，W4A4G4 让 FLUX.1-dev 12B 内存砍半、吞吐翻倍","fourtune-4bit-flux-12b","2026-07-09T00:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"52e96e5b-ec3b-4552-b94a-93cc2702ab84","Flow-Map GRPO：为确定性「少步生图」打开强化学习大门","flow-map-grpo-few-step","2026-07-05T12:02:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0e44f256-e66e-495c-82e3-aae4dd5e2374","LiveEdit 把扩散视频编辑推到 12.66 FPS：清华让 AR 实时编辑走出 PPT","liveedit-ar-video-editing","2026-07-01T06:15:00+00:00"]