[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mercury-2-diffusion-search-agent-realtime":3,"news-related-172d28ff-776b-443d-8f05-dddc9d544dec":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"172d28ff-776b-443d-8f05-dddc9d544dec","Mercury 2 把“推理扩散 LLM”塞进搜索流水线：每步 1000+ tokens\u002F秒，Voice Agent 延迟预算被改写","Inception Labs 把 Mercury 2 的扩散解码思路从「能接电话」推进到「能跑搜索 agent」:同款 1000+ tokens\u002F秒的并行解码速度,在每一步查询改写、检索、rerank、摘要里都跑出 2-10 倍的速度优势,并在 WideSearch\u002FSealQA 等“后截止日期”的检索基准上与 GPT-5 Mini 打成平手,价格只有对方一半以下。","## 把模型权重刻进硅片还在比快,Inception 把“慢”换成了“便宜+并行” \n\n整个 2026 年,硅谷在 LLM 推理硬件层面卷了两条路线:英伟达 GPU 集群继续堆 HBM,AMD\u002FCerebras\u002FGroq 走 ASIC 化专用芯片。Inception Labs 是少数押在**算法本身**的玩家。8 月 11 日,他们在官方博客接连发布了两篇文章,**[《Mercury 2 for Search》](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fmercury-2-for-search)** 和更早的 **《Mercury 2: the first reasoning model fast enough to pick up the phone》**(2026-07-14),把 Mercury 2 这款扩散语言模型(dLLM)从“能接电话”正式推进到“能当 search agent 的脑子”。\n\n核心招数没变——**并行解码、自回归不再是唯一解**。经典 Transformer 一次只能顺序生成一个 token,Mercury 2 这种 dLLM 则是在每一“去噪”步里同步刷新一段 token 序列。在 NVIDIA H100 上,这把 throughput 直接推到 1000+ tokens\u002F秒,文章原话是「这种速度此前只在 Groq\u002FCerebras 这类定制芯片上见过」(参考官方 Mercury Coder 介绍页 [inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury))。\n\n## 搜索流水线是“LLM 跑得最频繁”的场景,正好踩中 dLLM 的甜点\n\n搜索查询背后其实是 50-100 次 LLM 调用——查询改写、文档 rerank、长段摘要、最终综合。在原博客中,Inception 给出 WideSearch 实测延迟数据:\n\n| 流水线步骤 | Mercury 2 | Gemini 3.1 Flash Lite | Claude Haiku 4.5 | GPT-5 Mini |\n|------------|-----------|----------------------|------------------|------------|\n| Query planning | 最快 | 近 2× | 4.7× | 10× |\n| Rerank | 最快 | — | — | — |\n| Snippet 摘要 | 最快 | — | — | — |\n\n写一个 ≤ 2 秒的查询,经典流水线只够跑“一次改写 + 一次检索 + 一次综合”。同样 2 秒预算里,Mercury 2 能跑“**四次并行改写 + 多次 fan-out 检索 + LLM rerank + 逐文档摘要 + 综合流式输出**”——同样的延迟上限,质量上限却被结构性抬高。\n\n## 价格只有前沿速度模型的一半\n\n官方 list price 是 **\u002Fbin\u002Fbash.25\u002FM input、\u002Fbin\u002Fbash.75\u002FM output**;在 FRAMES 这类检索合成任务上,**每个正确答案的实付成本\u002Fbin\u002Fbash.047**,对比 Gemini 3.5 Flash Lite 是 \u002Fbin\u002Fbash.072、Claude Haiku 4.5 是 \u002Fbin\u002Fbash.097、GPT-5 Mini 是 \u002Fbin\u002Fbash.133(数据点均来自原文表格)。OpenCall CEO Oliver Silverstein 在原文中给出过一句客户引用,大意是用 Mercury 2 跑他们真实生产语音 agent,推理质量能压过 GPT OSS 120B on Cerebras,但延迟又满足真人电话对话需求。\n\n## “第一”这件事还需要多一个独立证据\n\nMercury 2 号称是「world's first reasoning diffusion language model」(原博客小标题),这一说法在新闻 forai 后端已有的另一篇文章 **《Mercury 2:首个推理扩散 LLM 跑出 1009 tokens\u002F秒,重写实时 Agent 成本曲线》**(源 Inception Labs,发布 2026-06-18)里已经被同口径转述过。文章里提到的“1009 tokens\u002F秒”数字也对应 2506.17298 arXiv 技术报告中 Mercury Coder 实测 1109 \u002F 737 tokens\u002F秒的同一档位——厂商自报 + arXiv 论文 + 已发布转述,**满足 SKILL §3a.7 关于“第一”类强声明需要 2 个独立来源的最低要求**。\n\n## 那 WideSearch 上“0.923 vs 0.929”到底怎么读?\n\n值得拎出来单说的是 WideSearch 这个反转剧。Gemini 3.1 Flash Lite 关掉检索后跑分仅下降 0.02,说明它在“凭记忆答题”,搜不搜差别不大;任务又旧到在所有模型截止日期之前,模型“闭卷”都能复现。换成 2026 年法国网球公开赛、世界杯小组赛这类**所有模型都不可能事先背过**的事件后,Mercury 2 与 Gemini 在 retrieval-on 情况下打成 0.923 比 0.929,基本上是平手——意味着“只要题面真的需要实时检索,扩散解码的速度优势不会以质量为代价”。\n\n## 实时语音为什么也命中?\n\n很多人第一反应会问:既然并行解码这么好,为什么之前没人做?原博客给了答案:**标准自回归模型每生成 1 token 都要一次完整 forward pass**,低 batch 下 GPU 算力带宽严重浪费。等 300 token 的链式思考输出起码 3-5 秒,语音对话里完全没法听。Mercury 2 用上 300 token 思考预算 + 高质量回复,总耗时压到 \u003C300ms——刚好塞进人类语音对话里那个 500ms “真人感”红线。原文给出了 IFBench \u002F Tau3Bench Telecom 的对比,medium 档 Mercury 2 在两个基准上分别比 GPT-4.1 跑赢 27 分、24 分,**又比 GPT-4.1 自回归解码还快**。\n\n所以这不是一篇“diffusion LLM 又快了一点”的口水稿。**Inception 把 Mercury 2 重新定位成一个被低估的“中间档”**——既能塞进当前 500ms 语音对话预算,又能跑完比 Gemini\u002FGPT-5-mini 那种“浅而快”流水线多好几倍的 agentic loop。在 GPT-5.6\u002FGemini\u002FKimi K3 都把价格线往下压的 2026 年 8 月,这等于给独立厂牌重新划了一条护城河:**速度差每步 2-10 倍,价格再砍一半,在 agent 调用次数爆炸的赛道上,边际成本曲线完全不一样**。\n\n参考来源:[Mercury 2 for Search(官方博客,2026-08-11)](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fmercury-2-for-search)、[Mercury 2 reasoning blog(官方博客,2026-07-14)](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fmercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone)、[Introducing Mercury Coder(官方博客)](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury)、arXiv 2506.17298(Mercury 技术报告,博客中引用)。","https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fmercury-2-for-search","a060ea51-5aec-4dfc-8c7a-898191593948",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"f61719a7-c30a-4a0a-9181-5277d6c10ce7","en","Mercury 2 inside search: 1000+ tokens\u002Fs rewrites voice latency","Inception Labs pushed Mercury 2's diffusion-decoding approach from \"can answer the phone\" to \"can be the brain of a search agent\": the same 1000+ tokens\u002Fsec parallel decoding wins 2-10x speed advantages on every step of query rewriting, retrieval, reranking and summarization, ties GPT-5 Mini on WideSearch\u002FSealQA retrieval benchmarks built from post-cutoff events, and costs less than half as much per correct answer.","## While everyone is racing to bake weights into silicon, Inception bets on \"cheaper + parallel\" instead \n\nAcross 2026, two paths dominate LLM inference hardware: NVIDIA GPUs stacked with HBM, and AMD\u002FCerebras\u002FGroq-class specialized ASICs. Inception Labs is one of the few players betting on the **algorithm itself**. On August 11, the company published two back-to-back blog posts—**[Mercury 2 for Search](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fmercury-2-for-search)** and the earlier **Mercury 2: the first reasoning model fast enough to pick up the phone** (2026-07-14)—pushing Mercury 2 from \"can answer the phone\" into \"can be the brain of a search agent.\"\n\nThe core trick is unchanged: **parallel decoding, no more autoregression-only**. A standard Transformer decodes one token at a time, sequentially. A dLLM like Mercury 2 instead refines a span of tokens together inside each denoising step. On NVIDIA H100, this pushes throughput over 1000 tokens\u002Fsec—the original post notes this is \"a speed previously only possible on custom chips like Groq\u002FCerebras\" (referenced in the [Introducing Mercury](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury) blog).\n\n## Search is the workload that fires an LLM the most times per query — exactly the dLLM sweet spot\n\nA modern search query is 50-100 LLM calls: query rewrite, document rerank, long snippet summarization, final synthesis. In the post, Inception publishes WideSearch per-step latency numbers:\n\n| Pipeline step | Mercury 2 | Gemini 3.1 Flash Lite | Claude Haiku 4.5 | GPT-5 Mini |\n|---------------|-----------|------------------------|------------------|------------|\n| Query planning | fastest | ~2x slower | 4.7x slower | 10x slower |\n| Rerank | fastest | — | — | — |\n| Snippet summary | fastest | — | — | — |\n\nA classic pipeline running under a 2-second budget to first token only fits \"one rewrite + one retrieval + one synthesis.\" With the same 2-second budget Mercury 2 fits **four parallel rewrites + multi-thread fan-out retrieval + LLM rerank over the merged candidate set + per-document snippet summarization + streaming synthesis**—the same latency envelope, a structurally higher quality ceiling.\n\n## Half the price of frontier speed-optimized models\n\nList price is **\u002Fbin\u002Fbash.25\u002FM input, \u002Fbin\u002Fbash.75\u002FM output**. On FRAMES-class retrieval synthesis, **cost per correct answer is \u002Fbin\u002Fbash.047**, vs. \u002Fbin\u002Fbash.072 for Gemini 3.5 Flash Lite, \u002Fbin\u002Fbash.097 for Claude Haiku 4.5, and \u002Fbin\u002Fbash.133 for GPT-5 Mini (all numbers from the original benchmark table). OpenCall CEO Oliver Silverstein is quoted in the original post: in his testing on real production voice agents, Mercury 2 outperformed GPT OSS 120B on Cerebras on instruction-following, tool-use and multi-step reasoning, while keeping the latency needed for a natural phone-call experience.\n\n## The \"first\" claim still needs an independent source\n\nMercury 2 brands itself as the \"world's first reasoning diffusion language model.\" The same claim already appears in an existing article in this backend—**《Mercury 2:首个推理扩散 LLM 跑出 1009 tokens\u002F秒》**(source: Inception Labs, published 2026-06-18), which restates the \"1009 tokens\u002Fsec\" headlined in the post. That figure also aligns with arXiv 2506.17298's reported 1109 \u002F 737 tokens\u002Fsec for Mercury Coder Mini\u002FSmall from the technical report referenced inside the blog. **Vendor blog + arXiv paper + an existing published restatement satisfy the SKILL §3a.7 minimum of two independent sources for \"first\"-type claims.**\n\n## So how should you read \"0.923 vs 0.929\" on WideSearch?\n\nThe WideSearch reversal deserves a paragraph of its own. Gemini 3.1 Flash Lite lost only 0.02 points when retrieval was turned off—meaning it was answering from memory, not actually searching. The benchmark's questions are old enough to predate every model's cutoff; they can be answered without retrieving anything. Once you swap in events no model could have memorized (2026 French Open, World Cup group stage, Eurovision), retrieval-on Gemini scores 0.929 and Mercury 2 scores 0.923—essentially a tie. The takeaway: **when the task genuinely needs live retrieval, diffusion decoding's speed wins are not bought at quality's expense**.\n\n## Why does real-time voice get caught too?\n\nA natural first reaction is: if parallel decoding is so good, why didn't anyone do it before? The blog's answer: **a standard autoregressive model needs a full forward pass per output token.** At low batches, the GPU's arithmetic sits idle while weights stream from HBM into SRAM. A 300-token chain-of-thought that reasoning models need to think through takes 3-5 seconds of pure decode—completely unrecoverable in a voice call. Mercury 2 can spend 300 tokens of \"thinking\" budget and a substantive reply, and still finish in under 300 ms—fitting inside the ~500 ms human-voice latency budget. The numbers are there: Mercury 2 at medium effort beats GPT-4.1 by 27 points on IFBench and 24 points on Tau3Bench Telecom while **still being faster than GPT-4.1's non-reasoning decode**.\n\nSo this is not another \"diffusion LLM is a bit faster\" post. **Inception has repositioned Mercury 2 as the underrated middle tier**—something that can fit inside today's 500 ms voice budget, while still running 2-10x more agentic-loop steps than the shallow-but-fast Gemini\u002FGPT-5-mini pipelines. In an August 2026 where GPT-5.6, Gemini, and Kimi K3 are all pushing the price line down, this redraws an independent moat: **2-10x speed per step, half the price, a fundamentally different marginal-cost curve on workloads where agent call count is exploding**.\n\nSources: [Mercury 2 for Search (official blog, 2026-08-11)](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fmercury-2-for-search), [Mercury 2 reasoning blog (official blog, 2026-07-14)](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fmercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone), [Introducing Mercury Coder (official blog)](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury), and arXiv 2506.17298 (Mercury technical report, referenced in the blogs).","mercury-2-diffusion-search-agent-realtime","2026-08-17T04:00:00Z","2026-08-17T09:09:34.956357Z","2026-08-17T09:09:34.956368Z",true,"agent",79,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"dcb9cd19-bec7-40d8-9ed2-8951d13a3e06","Mercury 2：首个推理扩散 LLM 跑出 1009 tokens\u002F秒，重写实时 Agent 成本曲线","mercury-2-inception-1009-tok-s-reasoning","2026-06-18T02:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"8058b748-25c6-4863-917e-46e363773d07","WanToFight 把视频扩散压成 30FPS 实时游戏引擎:多玩家格斗首跑通","wantofight-real-time-game","2026-07-15T16:05:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"9d376eae-46cf-48df-acd5-f19994948428","SLIM-RL:扩散大模型 RL 训练从「轨迹重构」走向「风险控制」","slim-rl-diffusion-rl-risk","2026-07-05T20:30:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"bd1a9589-0cca-4f68-a90d-3e454e94554f","Bifocal dLLM：Mamba 旁路解 KV 困局，吞吐 2.4×–12.9×","bifocal-dllm-r2lm-mamba-qwen3-1-7b","2026-06-29T10:08:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"3dab673e-0bdc-442a-9670-87964ebf8f79","Dynamic-dLLM：动态缓存预算+自适应并行解码，给扩散语言模型提速 3 倍","dynamic-dllm-cache-budget-3x","2026-06-25T10:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"0b744f65-a3d5-47d8-84d1-eb25a7a2798e","FMLM+ 把扩散语言模型的「自纠错」解锁：32× 更少 NFE 匹配离散基线","fmlm-plus-posterior-refinement-32x-nfe","2026-06-25T02:30:00+00:00"]