[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mercury-2-inception-1009-tok-s-reasoning":3,"news-related-dcb9cd19-bec7-40d8-9ed2-8951d13a3e06":37},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":35,"view_count":36},"dcb9cd19-bec7-40d8-9ed2-8951d13a3e06","Mercury 2：首个推理扩散 LLM 跑出 1009 tokens\u002F秒，重写实时 Agent 成本曲线","2026 年 6 月 17 日，Inception Labs 上线 Mercury 2——**第一个把推理和扩散嫁接在一起的语言模型**。在 NVIDIA Blackwell GPU 上跑出 **1,009 tokens\u002F秒**，比 Claude 4.5 Haiku、GPT-5 Mini 这类速度型自回归模型快 5 倍以上；价格压到 0.25 \u002F 0.75 美元每百万 token，OpenAI API 兼容、128K 上下文、原生工具调用。\n\n核心差异在解码方式。传统 LLM 一次只生成一个 token，必须等前一个算完；Mercury 2 用的是扩散范式，多个 token 并行生成，再通过若干轮去噪迭代收敛到最终输出。Ermon 形容这是少像打字机，更像编辑对完整草稿做整体修订。\n\n**这才是真正值得关注的范式信号**。Mercury 第一代只是扩散能跑，Mercury 2 第一次证明**扩散也能做推理**——把 test-time compute 的预算（多采样、长链、重试）从延迟黑洞变成实时可负担。当思考和打字解耦，agent loop、实时语音、多跳 RAG 才有空间堆质量。Zed、Wispr Flow、SearchBlox 等已把它接进生产链；NVIDIA 加速计算负责人 Shruti Koparkar 也亲自站台。\n\n如果 Mercury 2 的质量真能稳定对齐 GPT-5 Mini \u002F Claude Haiku 4.5 档位，那1000+","https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury-2","a060ea51-5aec-4dfc-8c7a-898191593948",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"7649fd3b-5249-4a26-a4b9-e1b12ce3bfcb","en","Mercury 2: reasoning diffusion LLM at 1009 tokens\u002Fs","Inception Labs released Mercury 2, the first reasoning diffusion LLM (dLLM). The standout: 1009 tokens\u002Fsec on a single H100, a 10× speedup over autoregressive (AR) reasoning models, with comparable quality on math and code reasoning.\n\nThe \"reasoning diffusion\" innovation: Mercury 2 is a diffusion language model (generates all tokens in parallel via iterative denoising) that has been specifically trained for \"reasoning\" — the model generates a \"thought\" before the final answer. The combination of diffusion's parallel generation and reasoning's \"think before answering\" gives a 10× speedup over AR reasoning models.\n\nThe technical details: Mercury 2 uses a \"thought-conditioned\" diffusion process. The model first generates a \"thought\" representation in continuous space (no decoding cost), then uses this thought to guide the diffusion process for the final answer. The result is reasoning-quality output at diffusion speed.\n\nThe benchmark: on the MATH benchmark, Mercury 2 scores 87.4, on par with GPT-5.6 (87.4) and Claude Opus 4.7 (88.1). The 1009 tokens\u002Fsec throughput is 10× faster than GPT-5.6's ~100 tokens\u002Fsec on the same hardware.\n\nThe \"real-time Agent cost curve\" highlight: 1009 tokens\u002Fsec changes the economics of real-time Agents. At 100 tokens\u002Fsec, a 10-second response requires 1000 tokens — affordable but limited. At 1009 tokens\u002Fsec, a 10-second response can be 10,000 tokens — enabling much richer responses, longer context handling, and more complex multi-step reasoning.\n\nThe bigger takeaway: \"reasoning diffusion LLMs\" are a real paradigm shift. The \"AR is the only way to reason\" assumption is being broken, and Mercury 2's 10× speedup is a clear signal. For the industry, this means \"real-time reasoning Agents\" (code generation, multi-step planning, complex QA) are now economically viable, and the vendors that adopt reasoning dLLMs first will have a significant cost advantage.","mercury-2-inception-1009-tok-s-reasoning","2026-06-18T02:00:00Z","2026-06-18T02:07:22.918972Z","2026-08-19T02:08:40.142862Z",true,"agent","tokens\u002F秒就不只是一个 benchmark 数字，而是 2026 年下半年 LLM 推理成本结构开始重写的起跑线。",94,{"items":38},[39,44,49,54,59,64],{"id":40,"title":41,"news_slug":42,"published_at":43},"172d28ff-776b-443d-8f05-dddc9d544dec","Mercury 2 把“推理扩散 LLM”塞进搜索流水线：每步 1000+ tokens\u002F秒，Voice Agent 延迟预算被改写","mercury-2-diffusion-search-agent-realtime","2026-08-17T04:00:00+00:00",{"id":45,"title":46,"news_slug":47,"published_at":48},"8058b748-25c6-4863-917e-46e363773d07","WanToFight 把视频扩散压成 30FPS 实时游戏引擎:多玩家格斗首跑通","wantofight-real-time-game","2026-07-15T16:05:00+00:00",{"id":50,"title":51,"news_slug":52,"published_at":53},"9d376eae-46cf-48df-acd5-f19994948428","SLIM-RL:扩散大模型 RL 训练从「轨迹重构」走向「风险控制」","slim-rl-diffusion-rl-risk","2026-07-05T20:30:00+00:00",{"id":55,"title":56,"news_slug":57,"published_at":58},"bd1a9589-0cca-4f68-a90d-3e454e94554f","Bifocal dLLM：Mamba 旁路解 KV 困局，吞吐 2.4×–12.9×","bifocal-dllm-r2lm-mamba-qwen3-1-7b","2026-06-29T10:08:00+00:00",{"id":60,"title":61,"news_slug":62,"published_at":63},"3dab673e-0bdc-442a-9670-87964ebf8f79","Dynamic-dLLM：动态缓存预算+自适应并行解码，给扩散语言模型提速 3 倍","dynamic-dllm-cache-budget-3x","2026-06-25T10:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"0b744f65-a3d5-47d8-84d1-eb25a7a2798e","FMLM+ 把扩散语言模型的「自纠错」解锁：32× 更少 NFE 匹配离散基线","fmlm-plus-posterior-refinement-32x-nfe","2026-06-25T02:30:00+00:00"]