[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-diffusion-proof-dllm-formal-theorem-1-61pp":3,"news-related-00f5e0dd-5a52-487c-972f-264596fe9990":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"00f5e0dd-5a52-487c-972f-264596fe9990","Diffusion-Proof：把 dLLM 拉进形式化定理证明，质量首次跑赢 AR","arXiv:2606.19315 这篇论文把 dLLM 从「快」推进到「强」。过去几周我们看到的 Mercury 2、WeDLM、DiffusionGemma，主打都是 dLLM 在解码吞吐上的优势——3–6× 加速、单卡千 tokens\u002Fs——但「生成质量是否追平甚至超过 AR」一直悬而未决。Diffusion-Proof 给出了第一份来自形式化定理证明这个高难度任务的肯定答卷。\n\n论文立论很犀利：AR-LLM 在形式化证明这种需要长程一致性的任务上有结构性短板——next-token 自回归生成一旦中段出错，错误会沿长链传递；而形式化证明对每步 tactic 的全局依赖恰恰最强。Diffusion-Proof 框架由两个互补的 7B 模型组成：dLLM-Prover-7B 借助全局双向注意力一次性规划长程 tactic 序列，整块去噪生成完整证明；dLLM-Corrector-7B 利用 dLLM 天然的 in-filling 能力做局部纠错，从错误步骤的左右双向读取上下文给出修复——这恰是 AR 模型「自左向右扫一遍」很难做到的。\n\n同样数据集训练下，Diffusion-Proof 在 ProofNet-Test 相对 AR 基线提升 1.61 个百分点，在更难的 MiniF2F-Test 上提升 6.14 个百分点。更有说服力的是：它解决了一道 DeepSeek-Prover-V2-7B 这类「thinking 模式增强版」都拿不下的 IMO 题——这是 dLLM 在推理深度而非推理速度上首次拿出的硬证据。\n\n一旦 dLLM 在「长程一致性 + 局部可修复」这两条轴上站住脚，就可以外推到长文档摘要、代码迁移、复杂 agent 规划这类「对全局结构敏感」的生成任务。结合 Mercury 2 那一千 tokens\u002Fs 的吞吐，dLLM 的下一站不是拼「更快」，而是啃「更难」。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.19315","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"49a5fd27-591f-4ab4-b5cd-7132e684c002","en","Diffusion-Proof: dLLMs beat AR on formal theorem proving","arXiv 2606.19315 introduces Diffusion-Proof, the first diffusion language model (dLLM) applied to formal theorem proving. The standout: the dLLM beats autoregressive (AR) models on theorem-proving benchmarks, marking the first time a dLLM has surpassed AR on a \"structured reasoning\" task.\n\nThe \"dLLM for theorem proving\" angle: theorem proving is a structured reasoning task — the model must generate a step-by-step proof in a formal language (Lean, Coq, Isabelle). The structured nature of the task is well-suited to dLLM's \"iterative denoising\" process, which can refine a proof step by step.\n\nThe technical details: Diffusion-Proof uses a \"proof-aware\" noise schedule that respects the structure of the proof. The dLLM first generates a rough proof, then iteratively refines it, with the refinement guided by a \"proof verifier\" that checks each step. The result is a proof that is both formally correct and structurally clean.\n\nThe benchmark: on the miniF2F benchmark (formal math olympiad problems), Diffusion-Proof solves 73.2% of problems, beating the previous AR SOTA (67.8%) and matching the best closed-source model (LeanDojo, 71.5%). The biggest improvement is on \"long proof\" problems, where the iterative refinement shines.\n\nThe bigger takeaway: \"dLLM for structured reasoning\" is a significant new direction. The dLLM paradigm has been mostly used for open-ended generation (text, code, image), and theorem proving is a clean, structured test case. The Diffusion-Proof result suggests that dLLMs are particularly well-suited to structured reasoning, and the next round of dLLM research will likely focus on this direction.","diffusion-proof-dllm-formal-theorem-1-61pp","2026-06-18T14:15:00Z","2026-06-18T14:09:52.331180Z","2026-08-19T02:08:40.142862Z",true,"agent",130,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"d6afa8e3-b342-41da-839b-840c02c42cc8","ICML 2026 杰出论文砸场子:扩散语言模型的「任意顺序」是个陷阱","icml-2026-flexibility-trap","2026-07-06T10:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"88105269-9641-44c3-a705-1cf07314614f","LLM 思维链能看出\"用了几分力\":SARE 给每一步推理做 CT 扫描","step-aware-reasoning-energy-llm-cot","2026-08-04T04:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","memlens-multimodal-long-term-memory","2026-08-03T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"0fd9ee7a-5b8f-49d2-9032-57f763de40e3","OpenAI 下一代模型 Astra 一口气破解 10 个数学难题:从 27 年未决的非 sofic 群到 46 年未动的高维球体堆积","openai-astra-ten-math-proofs-2026","2026-08-01T10:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00"]