[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-fmlm-plus-posterior-refinement-32x-nfe":3,"news-related-0b744f65-a3d5-47d8-84d1-eb25a7a2798e":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"0b744f65-a3d5-47d8-84d1-eb25a7a2798e","FMLM+ 把扩散语言模型的「自纠错」解锁：32× 更少 NFE 匹配离散基线","非自回归语言模型的两条主线——Masked Diffusion Models (MDM) 与 Flow Map Language Models (FMLM)——长期处于两难取舍中：MDM 灵活、支持任意顺序解码，但同时生成多个 token 时会出现分解误差，质量崩盘；FMLM 用联合序列传输绕开了这一瓶颈，单步就能产出像样的文本，却牺牲了推理时\"挑错重写\"的能力。\n\nCMU 与 KAIST 联合团队在 arXiv 2606.24773 上提出的 FMLM+ 与 Posterior Refinement，把这两条路径缝合到了同一个框架内。核心思路是给 FMLM 装上 masking 风格的噪声调度：一方面保留 FMLM\"一次性全局生成\"的特性，另一方面借来 MDM 的\"任意位置置信度评分\"，让模型在生成完整序列之后回过头检查每个 token 的全局一致性，按需重写。这一后验打分机制天然支持迭代自纠错——论文称之为 Posterior Refinement。\n\n实验结果显示，FMLM+ 仅需 32 轮 Posterior Refinement 就能逼近需要 1024 次函数评估的离散基线，最高做到 32× 更少的 NFE（Number of Function Evaluations）。覆盖 TinyStories、OpenWebText、GSM8K、Sudoku 四类基准，在所有数据集上一致压过 MDM 与 FMLM 同类方法。作者团队（CMU 的 Manan Agarwal、Sheel Shah 与 KAIST 的 Chanhyuk Lee、Jinwoo Kim 等）已经把代码与项目页放出，结构上兼容既有的大规模 MDM 预训练权重。\n\n这条工作的真正价值在于把\"快速生成\"和\"迭代修正\"放回到同一个可微框架内——这是离散 AR 模型结构性做不到的事。一旦这条路被打通，推理时 scaling、agentic self-edit、多轮 long-form 写作等高算力场景都多了一个不堆 GPU 也能提质的选项。扩散语言模型从\"论文里的并行玩具\"开始跨进\"工业可用\"的工程门槛，FMLM+ 至少贡献了一块砖。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.24773","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"7f32aad8-2945-4ce1-9550-6ac1c55fdc49","en","FMLM+ unlocks diffusion self-correction: 32x fewer NFEs","arXiv 2606.24773 introduces FMLM+ (Flow Matching Language Model Plus), a diffusion language model that achieves \"self-correction\" capability — a property previously considered exclusive to autoregressive (AR) models. The result: FMLM+ matches the quality of the discrete baseline with 32× fewer NFE (Number of Function Evaluations), the metric for diffusion model compute.\n\nThe \"self-correction\" property: AR models can naturally correct their own outputs — if a token is wrong, the model can revise it in a later step. Diffusion language models traditionally lack this — they generate all tokens \"in parallel\" and there's no natural way to \"go back and fix.\" FMLM+'s fix: a \"two-stage denoising\" process where the first stage generates a rough draft, and the second stage \"corrects\" the draft by re-masking low-confidence tokens and re-generating them. The \"correction\" is learned end-to-end, not hand-designed.\n\nThe result: FMLM+ matches the discrete (AR) baseline on MMLU and HumanEval with 32× fewer NFE. At the same NFE, FMLM+ is 2-3 quality points above the discrete baseline. The \"correction\" mechanism is general and can be added to any diffusion LLM.\n\nThe bigger takeaway: diffusion LLMs are catching up to AR on the things AR was thought to be uniquely good at. \"Self-correction\" was the last big advantage of AR, and FMLM+ closes that gap. The \"diffusion vs AR\" debate is increasingly irrelevant — both can match each other on quality and speed, and the choice will be driven by use case (parallel generation, controllable generation, etc.).\n\nFor the industry, this signals that diffusion LLMs are entering the \"production-ready\" phase. The next round of competition will be in \"diffusion LLM tooling\" — inference frameworks, fine-tuning recipes, and deployment guides.","fmlm-plus-posterior-refinement-32x-nfe","2026-06-25T02:30:00Z","2026-06-25T02:19:47.291680Z","2026-08-19T02:08:40.142862Z",true,"agent",116,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"bd1a9589-0cca-4f68-a90d-3e454e94554f","Bifocal dLLM：Mamba 旁路解 KV 困局，吞吐 2.4×–12.9×","bifocal-dllm-r2lm-mamba-qwen3-1-7b","2026-06-29T10:08:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"3dab673e-0bdc-442a-9670-87964ebf8f79","Dynamic-dLLM：动态缓存预算+自适应并行解码，给扩散语言模型提速 3 倍","dynamic-dllm-cache-budget-3x","2026-06-25T10:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"80bc5e25-d24d-45a1-9c2b-534ebcae39f9","腾讯 WeDLM 开源：让扩散 LLM 在标准因果注意力下跑出 3-6× vLLM 加速","tencent-wedlm-diffusion-llm-causal-3-6x","2026-06-16T20:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"634ba8f7-771e-4492-aacb-549112a8c91a","EPIC 让扩散语言模型重获并行优势：CFG 约束解码推理时间压缩 67.5%","epic-cfg-dllm-67-5pct-speedup","2026-06-09T16:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b45f5c46-c982-4410-9452-07a9f779218f","On-Policy Distillation 把扩散语言模型训练成本砍到 1\u002F15~1\u002F7000","opdlm-on-policy-distillation-dllm-1-7000x","2026-06-08T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"aa9841f2-87cd-48dc-a3cf-e0e47367b0af","UltraFlux：CVPR 2026稀疏注意力优化方案，4K上下文重建质量与效率双突破","ultraflux-cvpr-2026-4k-resonance-rope","2026-06-03T22:02:00+00:00"]