[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mercury-2-5-diffusion-llm":3,"topics-all":38,"news-related-ee700812-e4e5-4fa5-8687-6a0d7e5c7f78":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"ee700812-e4e5-4fa5-8687-6a0d7e5c7f78","Mercury 2.5 发布：扩散 LLM 跑出 1107 tokens\u002F秒，智能较上代提升 40%","Inception 发布 Mercury 2.5 扩散语言模型：智能较 Mercury 2 提升 40%，1107 tokens\u002F秒，260K 上下文，定价 $0.20\u002F$0.75 每百万 tokens。对比对象是 Luna\u002FFlash-Lite\u002FHaiku 等成本档模型，目前尚无独立 benchmark。","过去两年，几乎所有生产级 LLM 都在用同一种方式生成文本：自回归，一个 token 接一个 token。这条路径把成本和延迟直接绑在推理深度上——模型想得越多，回答越慢、越贵。Inception 走的是另一条路：扩散式语言模型（dLLM），先生成整段草稿，再并行精炼 token。9 月 8 日，这家公司发布了 [Mercury 2.5](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury-2-5)，官方称之为其迄今最强的生产级扩散模型。\n\n## 核心规格：智能 +40%，1107 tokens\u002F秒\n\n数字相当密集：智能水平较 Mercury 2 提升 40%，官方定位与 GPT-5.6 Luna (Low)、Gemini 3.5 Flash-Lite、Claude Haiku 4.5 等「成本优化型」前沿模型同档；在广泛可得的 NVIDIA GPU 上跑到 1107 tokens\u002F秒；上下文窗口扩到 260K tokens（[新闻稿](https:\u002F\u002Fsg.finance.yahoo.com\u002Fnews\u002Finception-launches-mercury-2-5-163000864.html)确认上一代为 128K）；定价每百万 tokens 输入 0.20 美元、输出 0.75 美元，首发期直接打到 0.04\u002F0.15 美元。能力侧补齐了可调推理、并行工具调用与 schema 对齐的 JSON 输出。官方还自称「据我们所知，这是迄今训练过的最大扩散语言模型」——厂商自述，不是第三方结论。\n\n## 生产场景先于跑分说话\n\nInception 的训练方法论值得一提：他们明说这次的 eval 是用客户反馈和生产失败案例打磨的，不是只盯 benchmark。两个落地数字值得记：AI 电话公司 OpenCall 称切换后 P99 响应从「数分钟」降到 1 秒、P50 从 0.4 秒降到 0.2 秒以内；编码工具厂商 Augment Code 把上下文压缩迁到 Mercury 后延迟降 82%（约 150 秒缩到 27 秒）、成本降 90% 且质量不变。同期开启预览的还有 Mercury Voice（首 token 延迟低于 170ms）与 Mercury Router（用 dLLM 做模型路由）。\n\n## 冷静的部分：还没有独立 benchmark\n\n意大利媒体 [tech-insider.org](https:\u002F\u002Ftech-insider.org\u002Fit\u002Fmercury-2-5-diffusion-llm-guida-api-2026) 的评测指南提醒：发布数天后 Mercury 2.5 仍无一份独立发布的 benchmark，1107 tokens\u002F秒是官方自测数字；正确姿势是拿自己的负载实测。X 上的讨论也在做算术：按 Artificial Analysis 9 月 8 日速度榜（榜首约 221 t\u002Fs）粗算，1107 tokens\u002F秒约为五倍——「若成立」三个字正是关键。\n\n所以呢：如果你的 Agent 里塞满高频短调用——查询改写、重排、压缩、路由——dLLM 已经从论文走进了选型表的正式一列；但在自己压测出数字之前，「最快」两个字请先放在引号里。","https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury-2-5","d34e492b-4ef4-4694-b6ad-fd4e58d0169a",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"da717eb3-1af3-4440-a779-b17f94ff9709","en","Mercury 2.5 Ships: 1,107 Tokens\u002Fsec, +40% Intelligence","Inception releases Mercury 2.5 diffusion LLM: +40% intelligence, 1,107 tokens\u002Fsec, 260K context, $0.20\u002F$0.75 per million tokens. No independent benchmarks yet.","For the past two years, nearly every production LLM has generated text the same way: autoregressively, one token at a time. That ties cost and latency directly to reasoning depth — the more a model thinks, the slower and pricier the answer. Inception took a different route: diffusion language models (dLLMs) draft a whole span of tokens, then refine it in parallel. On September 8 the company released [Mercury 2.5](https:\u002F\u002Fwww.inceptionlabs.ai\u002Fblog\u002Fintroducing-mercury-2-5), calling it its most capable production diffusion model yet.\n\n## The specs: +40% intelligence at 1,107 tokens\u002Fsec\n\nThe numbers are dense: a 40% intelligence jump over Mercury 2, positioned against cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5; 1,107 tokens per second on widely available NVIDIA GPUs; a context window extended to 260K tokens (the [press release](https:\u002F\u002Fsg.finance.yahoo.com\u002Fnews\u002Finception-launches-mercury-2-5-163000864.html) confirms the previous generation sat at 128K); and pricing of $0.20 per million input tokens \u002F $0.75 output, slashed to $0.04\u002F$0.15 at launch. Capability-wise it adds tunable reasoning, parallel tool calls, and schema-aligned JSON output. The company also calls it \"to our knowledge, the largest diffusion language model ever trained\" — a vendor claim, not a third-party verdict.\n\n## Production numbers before benchmarks\n\nInception's methodology deserves note: the company says this release's evals were sharpened with customer feedback and production failure cases, not benchmarks alone. Two deployment numbers stand out. OpenCall, which builds AI phone agents, reports P99 response time falling from several minutes to one second and P50 from 0.4s to under 0.2s after switching. Augment Code moved context compaction to Mercury and cut latency 82% (roughly 150 seconds down to 27) while reducing cost 90% with quality maintained. Previews of Mercury Voice (sub-170ms time-to-first-token) and Mercury Router (a dLLM that routes prompts across models) opened alongside.\n\n## The sober part: no independent benchmarks yet\n\nItalian outlet [tech-insider.org](https:\u002F\u002Ftech-insider.org\u002Fit\u002Fmercury-2-5-diffusion-llm-guida-api-2026) cautions that days after launch, no independently published benchmark existed for Mercury 2.5 — the 1,107 tokens\u002Fsec figure is measured on Inception's own infrastructure, and the right move is testing latency on your own workload. On X, observers did the arithmetic against Artificial Analysis' September 8 speed leaderboard (about 221 t\u002Fs at the top): if the number holds, it is roughly five times the leader — and \"if it holds\" is doing a lot of work in that sentence.\n\nSo what: if your agents are full of high-frequency short calls — query rewrites, reranking, compaction, routing — dLLMs have moved from papers into a permanent column of the model-selection spreadsheet. But until you benchmark it yourself, keep the word \"fastest\" in quotes.","mercury-2-5-diffusion-llm","2026-09-09T17:05:00Z","2026-09-09T17:06:35.453571Z","2026-09-09T17:06:35.453586Z",true,"agent",165,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"ee62535f-b897-437b-8674-02801632dadb","DeepSeek V4.1 非对称架构首发:读题 8B 答题 16B,KV 缓存砍到初代的 1\u002F437","deepseek-v4-1-flash-ced-kv-cache","2026-09-10T15:10:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"3d8b9b1a-e038-466f-9b6b-304f911e35a7","Kimi K3 开源三件套 MoonEP\u002FFlashKDA\u002FAgentEnv:Moonshot 把 2.8T MoE 训练栈完整交底","kimi-k3-moonep-flashkda-agentenv","2026-07-28T04:30:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"a16a4f36-cb00-4f14-a394-80d49075a323","上海AI Lab 发布 Intern-S2-Preview-397B：把「记忆」与「思考」拆开，397B 跑出万亿模型效果","shai-lab-intern-s2-preview-397b","2026-07-18T02:00:00+00:00"]