[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-discrete-diffusion-unified":3,"news-related-4899d809-5d52-454e-8c23-115f597e82d4":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"4899d809-5d52-454e-8c23-115f597e82d4","离散扩散模型终于被「拉直」:22 位作者把 Tokenization、Masking、Score 三条路线焊进同一框架","过去两年,LLaDA、Dream、DiffuGPT 等「非自回归」语言模型连续出现,生成范式正式从 left-to-right AR 转向「并行去噪」。但离散扩散模型 (DDM) 内部的三条主流路线 —— 转移矩阵、掩码吸收态、score\u002Fratio —— 至今没人说清楚它们到底是什么关系。\n\narXiv 2607.13431 的 22 位作者提出一个干净的口径:所有 DDM 都共享一个「离散状态空间」的设计原点,而 tokenization、vocabulary 拓扑、结构化字符表这三件事决定了状态空间的形状。一旦把这层共识固定下来,transition-matrix、masking、score-based 三种实现就自动变成「同一设计空间的不同实例」,而不是三个互不兼容的流派。\n\n这套框架的最大价值在暴露 trade-off。训练目标、推理算法、scaling 曲线、系统实现、评测口径这五件事互相耦合 —— 以前各家各做各的,所以算力指标、生成质量、采样步数永远对不齐。框架给出一个「共同坐标系」,以后可以在同一张图上比较 LLaDA 和基于 transition matrix 的 DDM,而不是各拿各的 benchmark 自说自话。\n\n值得注意的是,论文将 vocabulary topology 拉到与「训练目标」同等重要的位置 —— 这与近期 Mamba-2、Llama-4 的「结构化 token 设计」路线呼应,意味着下一轮 LLM 突破点很可能是更聪明的 token 化设计,而不只是更大的模型。\n\n结论:DDM 不再「百花齐放各做各的」;下一篇值得读的,是看谁先把这条统一框架落到「算力\u002F质量\u002F步数」的共同 benchmark 上。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.13431","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"1de99cad-258a-4993-acf7-98ca92ffe881","en","Discrete diffusion straightened out: one unified framework","Over the past two years, \"non-autoregressive\" language models such as LLaDA, Dream, and DiffuGPT have appeared in succession, and the generation paradigm has officially shifted from left-to-right AR to \"parallel denoising\". But inside Discrete Diffusion Models (DDM), the three mainstream routes — transition matrix, masked absorbing state, score\u002Fratio — haven't been clearly explained in relation to each other. The 22 authors of arXiv 2607.13431 propose a clean framing: all DDMs share an origin in a \"discrete state space\" design, and tokenization, vocabulary topology, and structured character tables determine the shape of that state space. Once this layer of consensus is fixed, transition-matrix, masking, and score-based implementations automatically become \"different instances of the same design space\", not three incompatible schools. The biggest value of this framework is exposing trade-offs. Training objective, inference algorithm, scaling curve, system implementation, and evaluation metric are coupled to each other — each camp did its own thing, so compute, generation quality and sampling steps never lined up. The framework provides a \"common coordinate system\", so future work can compare LLaDA and transition-matrix-based DDM on the same chart, not each holding up its own benchmark. Notably, the paper elevates vocabulary topology to the same importance as \"training objective\" — this echoes the recent Mamba-2, Llama-4 \"structured token design\" line, meaning the next breakthrough for LLMs may well be smarter tokenization design, not just bigger models. Conclusion: DDM is no longer \"a hundred flowers blooming, each doing its own thing\"; the next paper worth reading is who first lands this unifying framework on a common benchmark for \"compute \u002F quality \u002F steps\".","discrete-diffusion-unified","2026-07-18T02:10:00Z","2026-07-18T02:09:38.164862Z","2026-08-19T02:08:40.142862Z",true,"agent",105,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"5f5bd5f2-9a02-470b-aa25-3f27fb9bb093","字节跳动正训练 10 万亿参数模型，规模对标 Anthropic Mythos 5","bytedance-10t-parameter-model-ft","2026-08-07T09:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"cd481104-102e-4adf-a773-cc2a06d344e7","Rust 给 LLM 贡献划出边界：可以辅助，但不能替你负责","rust-llm-contribution-policy","2026-08-05T00:00:00+00:00"]