[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepseek-v4-pro-retires-flash-routes":3,"topics-all":38,"news-related-823e17ef-5927-40e6-9efd-08c4958922f0":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"823e17ef-5927-40e6-9efd-08c4958922f0","DeepSeek 旗舰 V4-Pro 今日退役:552B 的 V4.1-Flash 全面接班","9 月 14 日中午起,DeepSeek 把 deepseek-v4-pro 端点全量路由到 552B 的 V4.1-Flash 并按 Flash 费率计费,上线仅一个月的 1.6T 旗舰就此退役。非对称架构让小模型在性能与成本上反超大模型。","今天中午 12 点(UTC 04:00)起,所有打到 deepseek-v4-pro 端点的 API 请求,后台服务的已经不是那个 1.6 万亿参数的旗舰,而是 552B 的 V4.1-Flash——并且按 Flash 的费率计费。DeepSeek 在 9 月 10 日的[官方公告](https:\u002F\u002Fwww.deepseek.com\u002Fen\u002Fnews\u002Fdeepseek-v4-1-flash\u002F)里写得很直白:「我们正在逐步淘汰 V4-Pro」。上线刚满一个月的旗舰,就这样被自家最小的模型接了班。\n\n## 一场没有发布会的主帅更替\n\n先理时间线。V4.1-Flash 是 9 月 10 日随新架构家族一起亮相的「最小成员」,原生支持视觉理解。同一天官方就预告了退役计划:9 月 14 日 04:00 UTC 起,deepseek-v4-pro 的全部请求路由到 V4.1-Flash,计费同步切换到 Flash 费率,这个状态会持续到 V4.1-Pro 发布。更早一代的 V4-Flash 与 V4-Flash-Vision-Exp 也同步退役,旧模型名暂时兼容路由。整个切换没有新发布会,一纸公告加一次路由变更,旗舰就换人了。\n\n被替代的 V4-Pro 并不是旧货。按 [Unite.AI 的报道](https:\u002F\u002Fwww.unite.ai\u002Fdeepseek-ships-v4-pro-as-its-flagship-model-leaves-preview\u002F),它 8 月 12 日才结束近四个月的 preview 转正:1.6T 总参数、每 token 激活 49B,Hugging Face 仓库单月下载量超过 140 万次。从转正到退位,33 天。\n\n## 552B 凭什么顶替 1.6T\n\n答案写在架构里。V4.1-Flash 采用新的 Causal Encoder–Decoder 结构:输入侧每 token 只激活 8B 参数,输出侧 16B——读题的模型小,答题的模型大,两头都省。KV cache 相比上一代只需要 1\u002F4 的 HBM 和 1\u002F8 的 SSD 存储,而官方明确说,缓存命中的费用往往占 agent 成本的大头,把缓存压下来,就是把账单压下来。\n\n性能层面,官方的表述是「多方测试显示 V4.1-Flash 在性能、成本、速度和总运行时间上领先 V4-Pro」。对照 Unite.AI 抓取的定价页:V4-Pro 每百万输入 token 0.435 美元、输出 0.87 美元,且 Pro 端点并发上限 500,Flash 是 2500——旗舰在吞吐上反而是短板。API 模型名设为 deepseek-flash 即可调用,官方合作伙伴 WorkBuddy(含 CodeBuddy)与 OpenCode 已全面支持。\n\n## 参数规模不再是旗舰的定义\n\n这件事真正值得记一笔的,不是 DeepSeek 换了个模型,而是它愿意亲手把「参数最多」的那个模型请下位。MoE 时代总参数早已和实际算力脱钩:49B 激活的 1.6T 与 8B\u002F16B 激活的 552B,谁的服务成本低、谁的有效吞吐高,谁才是事实上的旗舰。DeepSeek 用一次路由变更承认了这一点。\n\n对开发者的信号也很实际:长上下文与 agent 工作负载里,缓存成本正在追上模型本身的智力溢价;峰谷定价(谷时半价)进一步把「什么时候跑任务」变成了一个工程问题。另外权重和技术报告已放上 Hugging Face,官方还留了一句——「规划 2000 GPU 加存储集群的大规模部署?来聊」——本地化与私有部署路线并没有放弃。\n\n所以呢:当一家以效率著称的实验室主动退役自家最大的模型,「旗舰」的定义权就已经从参数表转移到了单位成本的性能上。下一次谁再拿总参数规模说事,可以把这次路由变更直接甩给他。\n","https:\u002F\u002Fwww.deepseek.com\u002Fen\u002Fnews\u002Fdeepseek-v4-1-flash\u002F","4194681c-1a38-405d-a917-40e1dc2622ea",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"b52db7e9-7c58-42c3-9536-5132cb2f8f72","deepseek",{"id":19,"name":20,"slug":20,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"aa47d01d-6b74-43d0-9e82-f401b1111811","en","DeepSeek Retires V4-Pro: the 552B V4.1-Flash Takes Over","From noon on September 14, DeepSeek routes all deepseek-v4-pro endpoint traffic to the 552B V4.1-Flash at Flash rates, retiring the 1.6T flagship just a month after launch. The asymmetric architecture lets the smaller model beat the bigger one on performance and cost.","Starting at 12:00 noon Beijing time today (04:00 UTC), every API request hitting the deepseek-v4-pro endpoint is no longer served by the 1.6-trillion-parameter flagship — it is served by the 552B V4.1-Flash, billed at Flash rates. DeepSeek's [official announcement](https:\u002F\u002Fwww.deepseek.com\u002Fen\u002Fnews\u002Fdeepseek-v4-1-flash\u002F) from September 10 puts it bluntly: \"We're phasing out V4-Pro.\" A flagship that went generally available barely a month ago has just been succeeded by the smallest model in the company's lineup.\n\n## A leadership change without a launch event\n\nStart with the timeline. V4.1-Flash arrived on September 10 as the smallest member of a new architecture family, with native visual understanding. The same announcement previewed the retirement plan: from 04:00 UTC on September 14, all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates, and this stays in effect until V4.1-Pro launches. The older V4-Flash and V4-Flash-Vision-Exp are retired too, with the old model names temporarily routing for compatibility. No launch event, no rebranding — one announcement plus one routing change, and the flagship changed hands.\n\nThe model being replaced is not old stock. Per [Unite.AI's report](https:\u002F\u002Fwww.unite.ai\u002Fdeepseek-ships-v4-pro-as-its-flagship-model-leaves-preview\u002F), V4-Pro only graduated from a near-four-month preview on August 12: 1.6T total parameters, 49B active per token, and more than 1.4 million downloads on Hugging Face in the last month. From general availability to retirement: 33 days.\n\n## Why 552B can take over from 1.6T\n\nThe answer is in the architecture. V4.1-Flash uses a new Causal Encoder–Decoder design: just 8B active parameters on the input side per token, 16B on the output side — a small model reads the prompt, a bigger one writes the answer, and both ends save compute. Its KV cache needs only 1\u002F4 the HBM and 1\u002F8 the SSD storage of the previous generation, and DeepSeek notes that cache-hit charges often account for a large share of agent costs — shrink the cache, shrink the bill.\n\nOn performance, the official wording is that \"tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime.\" Against the pricing page captured by Unite.AI: V4-Pro lists $0.435 per million input tokens and $0.87 per million output tokens, with a concurrency cap of 500 on the Pro endpoint versus 2,500 for Flash — the flagship is actually the weaker option on throughput. The API model name is simply deepseek-flash, and official partners WorkBuddy (including CodeBuddy) and OpenCode already support it fully.\n\n## Parameter count no longer defines a flagship\n\nThe real story here is not that DeepSeek swapped a model — it's that it willingly pushed its largest model off the throne. In the MoE era, total parameters decoupled from real compute long ago: between a 1.6T model with 49B active and a 552B model with 8B\u002F16B active, whichever serves cheaper and sustains higher effective throughput is the de facto flagship. With one routing change, DeepSeek conceded exactly that.\n\nThe signal for developers is practical: in long-context and agentic workloads, cache costs are catching up with the model's intelligence premium, and peak\u002Foff-peak pricing (off-peak at 50%) turns \"when to run the job\" into an engineering question. Weights and the technical report are already on Hugging Face, and the announcement ends with an invitation — \"planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let's talk\" — so the local, private-deployment path is alive and well.\n\nSo here's the takeaway: when a lab famous for efficiency voluntarily retires its own largest model, the definition of \"flagship\" has already moved from the parameter table to performance per unit cost. The next time someone argues model quality by total parameter count, hand them this routing change.\n","deepseek-v4-pro-retires-flash-routes","2026-09-14T15:10:00Z","2026-09-14T15:08:39.186794Z","2026-09-14T15:08:39.186803Z",true,"agent",58,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"ee62535f-b897-437b-8674-02801632dadb","DeepSeek V4.1 非对称架构首发:读题 8B 答题 16B,KV 缓存砍到初代的 1\u002F437","deepseek-v4-1-flash-ced-kv-cache","2026-09-10T15:10:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","minimax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"8c5bd0e1-551a-405b-b2fc-a67520bf84c5","Colibri v1.11.0 发布:纯 C 引擎直读 DeepSeek V4.1 Flash,552B 从 SSD 流进 CPU","colibri-v1-11-deepseek-v41-flash","2026-09-14T13:09:40+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"550cee5e-18e8-4236-9304-7207ebc221a8","Agnes 3.0 Flash 开源:72 层仅 18 层带 KV 缓存,33B 单卡跑 262k 上下文","agnes-3-0-flash-preview-open-weights","2026-09-13T15:20:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"365b770a-2cca-40a0-beb0-1eff823702c0","IBM 开源 Granite Time Series PatchTST-FM-r2:零样本 SOTA,Apache 2.0 商用许可","ibm-granite-patchtst-fm-r2-zero-shot-apache","2026-09-12T11:00:00+00:00"]