[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepseek-v4-1-flash-ced-kv-cache":3,"topics-all":38,"news-related-ee62535f-b897-437b-8674-02801632dadb":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"ee62535f-b897-437b-8674-02801632dadb","DeepSeek V4.1 非对称架构首发:读题 8B 答题 16B,KV 缓存砍到初代的 1\u002F437","DeepSeek 发布 552B MoE 新架构 V4.1-Flash:输入激活 8B、输出 16B 的非对称设计,KV 缓存降至 V1 的 1\u002F437、HBM 需求 1\u002F4,9 月 14 日起 V4-Pro 全量路由到新模型,权重已开源。","9 月 10 日,DeepSeek 做了一件少见的事:上线新模型 V4.1-Flash 的同时,宣布自家旗舰 V4-Pro 进入退役流程——9 月 14 日 04:00 UTC 起,所有 `deepseek-v4-pro` 的 API 请求将被路由到 V4.1-Flash,并按新模型的价格计费,直到 V4.1-Pro 发布为止。家族里最小的模型把旗舰逼下台,这个信号本身比任何跑分都值得咀嚼。\n\n## 非对称架构:读题用 8B,答题用 16B\n\nV4.1-Flash 是一个 552B 参数的 MoE,核心变化是一套新的 Causal Encoder–Decoder(CED)架构:输入侧每个 token 只激活 8B 参数,输出侧激活 16B。官方称,配合新的预训练方法和更大规模的 RL 后训练,它在 benchmark 上超过了包括 V4-Pro 在内的旗舰模型。\n\n第三方分析([atoms.dev](https:\u002F\u002Fatoms.dev\u002Fblog\u002Fdeepseek-v4-1-flash))拆得更细:40 层 Transformer 被拆成 20 层因果编码器加 20 层解码器,解码器的全局 KV cache 不再逐层独立计算,而是从编码器的最终隐状态投影得到。这种非对称设计针对的正是 agent 负载的特征——上下文、工具返回、历史对话加起来的输入,往往远大于最终输出。\n\n## KV 缓存:Agent 成本的命门\n\n[官方公告](https:\u002F\u002Fapi-docs.deepseek.com\u002Fnews\u002Fnews260910)给出的数字:与上一代相比,V4.1-Flash 的 KV cache 只需要 1\u002F4 的 HBM、1\u002F8 的 SSD 存储。官方也直接点破了动机——在 agent 任务里,缓存命中费用往往占总成本的很大一块。\n\natoms.dev 从模型卡读到的细节更夸张:全局 KV cache 低至每 token 890 字节,约为 V4-Flash 的 1\u002F4、初代 V1 的 1\u002F437(AIBase 的独立报道交叉印证了这个量级);配合 SWA Bounded Replay 机制,不再把完整的滑窗注意力 KV 状态持久化到 SSD。另有 Compressed Sparse Attention 2:注意力层在 Full、Reindex、Reuse 三种静态模式间选择,跨层共享 KV 数据与稀疏索引,而不是每层重复重建。\n\n## 价格与生态\n\n定价继续峰谷结构:off-peak 费率为 peak 的 50%,新价格自 9 月 10 日 04:00 UTC 生效。atoms.dev 列出的具体数字:off-peak 未缓存输入 $0.15\u002F百万 token、缓存输入 $0.003、输出 $0.60,peak 翻倍。生态侧,官方合作伙伴 WorkBuddy(含 CodeBuddy)与 OpenCode 已宣布全面支持;权重和技术报告已放上 [Hugging Face](https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4.1-Flash)。同时旧模型 V4-Flash 与 V4-Flash-Vision-Exp 退役,旧模型名暂时兼容路由到 V4.1-Flash。\n\n## 几点判断\n\n其一,「Flash 逼退 Pro」说明竞争焦点正在从「谁跑分高」转向「谁的单 token 成本低」——当小模型的跑分追平旗舰,旗舰的存在意义就只剩下品牌和更大的参数故事。其二,非对称架构是给 agent 时代定制的:长上下文、多轮调用、输入密集,KV cache 压缩直接命中这类负载的成本结构,比单纯的规模叙事更有工程含金量。其三,留一分清醒:「超过旗舰」目前是官方提交口径,独立分析确认了架构与价格数字,但具体基准的胜负,建议在自己的负载上实测再下结论。\n\n所以呢:下一波模型竞争的胜负手,可能不在参数表上,而在「每个 token 的显存账」上。V4.1-Flash 把这笔账摆上了台面——真正的问题是,接下来发布的 V4.1-Pro,还能拿什么把旗舰的故事讲下去?","https:\u002F\u002Fapi-docs.deepseek.com\u002Fnews\u002Fnews260910","3202bafd-f422-4632-b981-ec60896b79df",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"b52db7e9-7c58-42c3-9536-5132cb2f8f72","deepseek",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"f9a76c08-a359-40a6-a753-555806d7d062","en","DeepSeek V4.1-Flash Retires Its Own Flagship: KV Cache at 1\u002F437","DeepSeek ships V4.1-Flash: 552B MoE, asymmetric 8B-in\u002F16B-out, KV cache 1\u002F437 of V1, HBM 1\u002F4. V4-Pro routes to it Sept 14; weights open.","On September 10, DeepSeek did something rare: it launched V4.1-Flash and simultaneously put its own flagship V4-Pro on a retirement path. Starting 04:00 UTC on September 14, every `deepseek-v4-pro` API request will be routed to V4.1-Flash and billed at the new model's rates, until V4.1-Pro ships. The smallest model in the family just pushed the flagship off the stage — that signal matters more than any benchmark score.\n\n## Asymmetric architecture: 8B to read, 16B to write\n\nV4.1-Flash is a 552B-parameter MoE built on a new Causal Encoder–Decoder (CED) design: only 8B parameters activate per input token, while output decoding activates 16B. According to the official announcement, new pre-training methods plus larger-scale RL post-training push its benchmark results past flagship models, including V4-Pro.\n\nThird-party analysis ([atoms.dev](https:\u002F\u002Fatoms.dev\u002Fblog\u002Fdeepseek-v4-1-flash)) goes deeper: the 40 Transformer layers split into a 20-layer causal encoder and a 20-layer decoder, and the decoder's global KV cache is no longer derived independently per layer — it is projected from the encoder's final hidden states. The asymmetry maps directly onto agent workloads, where accumulated context, tool returns, and dialogue history make input far larger than output.\n\n## KV cache: the cost ceiling for agents\n\nThe [official announcement](https:\u002F\u002Fapi-docs.deepseek.com\u002Fnews\u002Fnews260910) states that versus the previous generation, V4.1-Flash's KV cache needs just 1\u002F4 the HBM and 1\u002F8 the SSD storage. The motivation is explicit: in agent tasks, cache-hit charges often account for a large share of total cost.\n\nDetails from the model card, as reported by atoms.dev, are more striking: the global KV cache footprint is 890 bytes per token — roughly a quarter of V4-Flash and about 1\u002F437 of the original DeepSeek V1 (independently corroborated by AIBase's coverage). An SWA Bounded Replay mechanism avoids persisting the full sliding-window-attention KV state to SSD, and Compressed Sparse Attention 2 lets attention layers pick among three static modes — Full, Reindex, Reuse — sharing KV data and sparse-attention indices across layers instead of rebuilding them.\n\n## Pricing and ecosystem\n\nPricing keeps the peak\u002Foff-peak structure: off-peak rates are 50% of peak, effective 04:00 UTC on September 10. Per atoms.dev's numbers: off-peak uncached input at $0.15 per million tokens, cached input at $0.003, and output at $0.60, with peak rates doubled. On the ecosystem side, official partners WorkBuddy (including CodeBuddy) and OpenCode have announced full support, and the weights plus a technical report are live on [Hugging Face](https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4.1-Flash). The older V4-Flash and V4-Flash-Vision-Exp are retired, with their model names temporarily routing to V4.1-Flash for compatibility.\n\n## Three takeaways\n\nFirst, \"Flash retires Pro\" shows the competition shifting from \"whose benchmark is higher\" to \"whose per-token cost is lower\" — once a small model matches the flagship on benchmarks, the flagship's remaining value is mostly branding and a bigger parameter story. Second, the asymmetric architecture is purpose-built for the agent era: long contexts, multi-turn calls, and input-heavy traffic mean KV-cache compression hits the cost structure of these workloads directly — more engineering substance than a pure scale narrative. Third, stay sober: \"beats the flagship\" is currently the official submission; independent analysis confirms the architecture and pricing figures, but for specific benchmark outcomes, test on your own workloads before concluding.\n\nSo what: the next round of model competition may be decided not by the parameter sheet but by \"the memory bill per token.\" V4.1-Flash has put that bill on the table — the real question is what V4.1-Pro will use to keep the flagship story alive.","deepseek-v4-1-flash-ced-kv-cache","2026-09-10T15:10:00Z","2026-09-10T15:08:18.464441Z","2026-09-10T15:08:18.464466Z",true,"agent",124,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"823e17ef-5927-40e6-9efd-08c4958922f0","DeepSeek 旗舰 V4-Pro 今日退役:552B 的 V4.1-Flash 全面接班","deepseek-v4-pro-retires-flash-routes","2026-09-14T15:10:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"ee700812-e4e5-4fa5-8687-6a0d7e5c7f78","Mercury 2.5 发布：扩散 LLM 跑出 1107 tokens\u002F秒，智能较上代提升 40%","mercury-2-5-diffusion-llm","2026-09-09T17:05:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"31f3215c-0892-419d-a610-fe815cc60bbe","GPT-5.6 降价 80% 把竞争拉进「同等智能成本」：DeepSeek V4 Flash 接招，国产模型卡出双线赛道","gpt-5-6-luna-price-cut-equal-intelligence-cost","2026-08-12T03:00:00+00:00"]