[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-diffusion-llm-dllm-block-causal-flexible-iteration":3,"topics-all":36,"news-related-1bc27b84-63d1-4923-bfb1-0492782f0f9f":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1bc27b84-63d1-4923-bfb1-0492782f0f9f","扩散语言模型崛起：同一模型极速与高准确率兼得","传统自回归模型面临精度与速度二选一的困境——大模型推理慢，小模型不够准。扩散语言模型（dLLM）正在打破这一僵局。\n\ndLLM的核心思路是先生成一段包含占位符的粗略文本，再用双向注意力机制对整段文字迭代精炼。每次迭代让输出更准确，迭代次数越多质量越高。这意味着运行时可以在延迟和精度之间动态切换：语音助手需要毫秒级响应？用2-3步。复杂代码推理需要高质量？用20步以上。同一模型，无需维护多个版本或复杂路由逻辑。\n\n架构经历了三个阶段快速演进。第一代全上下文并行精炼但无法使用KV缓存，计算代价过高；第二代引入block-wise causal attention，以8-64 token为块进行局部精炼，开始具备实用价值；第三代持续优化token editing和流式解码等能力。\n\n对推理服务商和边缘设备而言，dLLM意味着更灵活的算力分配策略。开源社区已推出LLaDA 2.0-mini等轻量版本，可在消费级GPU上运行。当然，dLLM目前仍处于早期阶段，迭代精炼带来的额外延迟能否换来足够的精度收益，还需更广泛验证。但当模型架构本身开始打破大而慢、小而快的二元对立，AI部署的效率曲线将迎来显著改变。","https:\u002F\u002Fdevelopers.redhat.com\u002Farticles\u002F2026\u002F04\u002F28\u002Fbeyond-next-token-why-diffusion-llms-are-changing-game","552102e7-842e-4d02-baad-91df815abca5",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"5db78e9a-78cb-4aeb-a53b-aa65510b8178","en","Diffusion LLMs rise: speed and accuracy from one model","Traditional autoregressive models face a precision-vs-speed either-or dilemma — large models are slow, small models aren't accurate enough. Diffusion Language Models (dLLM) are breaking this deadlock.\n\nThe core idea of dLLM is to first generate a coarse text with placeholders, then use bidirectional attention to iteratively refine the entire text segment. Each iteration makes the output more accurate; the more iterations, the higher the quality. This means runtime can dynamically switch between latency and precision: voice assistant needs millisecond-level response? Use 2-3 steps. Complex code reasoning needs high quality? Use 20+ steps. Same model, no need to maintain multiple versions or complex routing logic.\n\nThe architecture has gone through three stages of rapid evolution. The first generation refined full context in parallel but couldn't use KV cache, with too high compute cost; the second generation introduced block-wise causal attention, doing local refinement in 8-64 token blocks, beginning to have practical value; the third generation continues optimizing token editing and streaming decoding capabilities.\n\nFor inference service providers and edge devices, dLLM means more flexible compute allocation strategies. The open-source community has released lightweight versions like LLaDA 2.0-mini, runnable on consumer-grade GPUs. Of course, dLLM is still in early stages; whether the extra latency from iterative refinement can be exchanged for sufficient precision gains needs broader validation. But when model architecture itself starts to break the binary opposition of big-and-slow vs small-and-fast, the efficiency curve of AI deployment will see significant change.","diffusion-llm-dllm-block-causal-flexible-iteration","2026-04-30T04:10:00Z","2026-04-30T04:07:47.152780Z","2026-08-19T02:08:40.142862Z",true,"agent",154,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"8914cfad-53bc-4f76-bb06-8579d23cfa5d","AR 大模型外挂 diffusion 权重:Uno 免草稿模型实现 3 倍无损解码加速","uno-diffusion-augmented-llm-speedup","2026-09-08T13:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"bd1a9589-0cca-4f68-a90d-3e454e94554f","Bifocal dLLM：Mamba 旁路解 KV 困局，吞吐 2.4×–12.9×","bifocal-dllm-r2lm-mamba-qwen3-1-7b","2026-06-29T10:08:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"3dab673e-0bdc-442a-9670-87964ebf8f79","Dynamic-dLLM：动态缓存预算+自适应并行解码，给扩散语言模型提速 3 倍","dynamic-dllm-cache-budget-3x","2026-06-25T10:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"0b744f65-a3d5-47d8-84d1-eb25a7a2798e","FMLM+ 把扩散语言模型的「自纠错」解锁：32× 更少 NFE 匹配离散基线","fmlm-plus-posterior-refinement-32x-nfe","2026-06-25T02:30:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"80bc5e25-d24d-45a1-9c2b-534ebcae39f9","腾讯 WeDLM 开源：让扩散 LLM 在标准因果注意力下跑出 3-6× vLLM 加速","tencent-wedlm-diffusion-llm-causal-3-6x","2026-06-16T20:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"634ba8f7-771e-4492-aacb-549112a8c91a","EPIC 让扩散语言模型重获并行优势：CFG 约束解码推理时间压缩 67.5%","epic-cfg-dllm-67-5pct-speedup","2026-06-09T16:00:00+00:00"]