[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepseek-v4-pro-0813-ga-1m-context-moe":3,"news-related-d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","DeepSeek V4 Pro 0813 8 月 12 日从 OpenRouter GA 上线,1.6T 总参 \u002F 49B 激活的 MoE,1M token 上下文窗口最大输出 384K,API 定价 0.435 美元\u002F百万输入、0.87 美元\u002F百万输出。Pro 端点并发上限 500,远低于 Flash 的 2500。","## 技术背景\n\nDeepSeek 在 2026 年 4 月 24 日以 MIT 协议同时开放了 V4 Pro 与 V4 Flash 的权重与 API,从这一天起就进入了\"先预览、再 GA\"的两段式发布。7 月 31 日 Flash 先转正,Pro 留 preview。4 个月之后,Pro 端点跟着转正,挂在 build 号 **0813** 上。\n\nV4 Pro 不是一个普通的 dense 升级,而是 1.6T 总参数、每 token 激活 49B 的 Mixture-of-Experts。架构上它把两种新注意力同时塞进了同一套模型:DeepSeek 称之为 **Compressed Sparse Attention (CSA)** 与 **Heavily Compressed Attention (HCA)**,官方给出的账本是:在 1M token 设定下,单 token 推理算力压到上一代 V3.2 的 27%,KV cache 压到 10%。两个 V4 模型都用了 32T+ token 做预训练,后训练则走\"按领域独立长专家 → on-policy 蒸馏合并回单一模型\"的路线。\n\nHugging Face 上 V4 Pro 的 preview 仓库过去一个月已经有 1.4M+ 次下载,官方建议本地跑最大推理档位时至少留 384K token 的上下文。\n\n## 核心内容\n\n**1) API 经济沿用 preview 价格**\n\n- 输入:cache miss **$0.435 \u002F 百万 token**,cache hit **$0.003625 \u002F 百万 token**\n- 输出:**$0.87 \u002F 百万 token**\n- 上下文窗口:**1,048,576 token**(1M),最大输出 **384,000 token**\n- Pro 端点并发上限 **500**,Flash 端点 **2,500**——Pro 仍然走\"慢但重\"的路线\n\nOpenRouter 上,deepseek-v4-pro-0813 页面标\"Released Aug 12, 2026\",与 DeepSeek 自家 API 文档里的 deepseek-v4-pro 版本号对齐。\n\n**2) 三个运行档位**\n\nDeepSeek 把 API 切成三档:\n- non-thinking\n- high reasoning effort\n- max effort(官方原文:\"pushing the boundary of model reasoning capability\")\n\nAPI 同时兼容 OpenAI ChatCompletions、Anthropic Messages 以及 DeepSeek 自家 Responses API,Pro 与 Flash 都支持 tool calling 与 JSON 输出。\n\n**3) 官方 benchmark(V4-Pro-Max 档位,均为厂商自报)**\n\n- SWE-bench Verified: **80.6%** resolved\n- Terminal Bench 2.0: **67.9%** accuracy\n- GPQA Diamond: **90.1%** pass@1\n- Humanity's Last Exam: **37.7%** pass@1\n- MMLU-Pro: **87.5%**\n- LiveCodeBench: **93.5%** pass@1\n- Codeforces rating: **3,206**\n- MRCR @ 1M tokens: **83.5 MMR**\n\n官方对比表里,V4-Pro-Max 在 Terminal Bench 2.0(67.9 vs 75.1)与 Humanity's Last Exam(37.7 vs 44.4)上仍落后 GPT-5.4 xHigh 与 Gemini-3.1-Pro;在 SWE-bench Verified(80.6)上与 Gemini-3.1-Pro 持平,略低于 Claude Opus 4.6(80.8);但在 LiveCodeBench 与 Apex Shortlist 上拿下表格最高分。**这些分数尚未有独立第三方对 0813 build 做完整复现。**\n\n**4) 商业侧:Pro 端价格会涨**\n\nDeepSeek 的 pricing 页面挂了一条提示:整体 API 定价在\"不久的将来会有显著上调\",具体以官方公告为准。在此之前,8 月 12 日的列表价继续生效。OpenRouter 上\"用户实际支付价\"明显低于 $0.435 的列表价,差异被解释为 cache 命中与折扣。\n\n**5) 还没出 0813 的开源权重**\n\nHugging Face 仓库目前只挂着 4 月的 preview build,DeepSeek 尚未公布 0813 权重的发布时间表,也未明示 GA build 相对 preview 是否只有后训练差异。V4 线的官方节奏写得很清楚:**先 API,后权重**。\n\n## 个人评论与行业影响\n\nV4 Pro 的 GA 不是技术奇袭,而是 DeepSeek 这家公司在公开做一次\"产品定位\"。\n\n第一, **API 先行 + 权重滞后** 已经是 DeepSeek V3 以来就稳定使用的发布节奏。把\"preview\" 这个标签摘掉的动作,只是把定价、并发、运行档位等商业参数固定下来,模型权重再分批放出——这是一种典型的\"先把生态位占住,再让社区做移植与基准复现\"的打法。\n\n第二, **CSA + HCA 的双注意力组合 + 27% \u002F 10% 的算力与 KV 账本** 是 DeepSeek 把\"长上下文\"重新拉回 MoE 阵营的关键。V3.2 时代 MoE 在 100K+ 上下文里 KV 增长过快,被许多团队放弃长文档;V4 Pro 用结构上的稀疏 + 重压缩,等于在 MoE 主线之外又开了一条\"1M token 仍然跑得动\"的次线。配合 OpenAI \u002F Anthropic 这一档同期把 GPT-5.6 \u002F Claude Opus 系列推上同级别 1M 上下文,基本可以下结论:**2026 下半年,1M token 是头部模型的入场券,而不是卖点。**\n\n第三, **\"将显著上调 API 价格\"** 的官方提示,意味着 DeepSeek 在 5000 亿元估值融资(报道见 36氪)之后,开始把\"靠低价抢市场\"切换成\"靠能力卖 API\"。配合 V4 Flash 0731 在 Artificial Analysis Intelligence Index 上以 50 分打 GPT-5.6 Luna 的 1 分差距,Pro 的提价空间在于\"V4-Pro-Max 在 SWE-bench \u002F LiveCodeBench 这类开发者工作流上已经追平甚至超过 GPT-5.4 \u002F Claude Opus 4.6\"——而开发者工作流对 token 单价的容忍度,远高于聊天\u002F摘要场景。\n\n留给读者的问题:**V4 Pro 0813 的开源权重一天不放出来,DeepSeek 一天就是在用\"preview → GA\"的措辞差做商业缓冲;一旦权重跟 GA 同步放出,MoE 长上下文赛道的下一波竞争会立刻被 Kimi、Qwen 与字节系的开源旗舰接过去。** 这条时间窗口,大约就是接下来 30 天。\n\n---\n\n**素材来源**:\n- Unite.AI(2026-08-12):DeepSeek Ships V4 Pro as Its Flagship Model Leaves Preview — https:\u002F\u002Fwww.unite.ai\u002Fdeepseek-ships-v4-pro-as-its-flagship-model-leaves-preview\u002F\n- OpenRouter Model Page(2026-08-12):DeepSeek V4 Pro 0813 — API Pricing & Benchmarks — https:\u002F\u002Fopenrouter.ai\u002Fdeepseek\u002Fdeepseek-v4-pro-0813\n- Hugging Face(2026-04 起,持续更新):deepseek-ai\u002FDeepSeek-V4-Pro model card — https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4-Pro","https:\u002F\u002Fwww.unite.ai\u002Fdeepseek-ships-v4-pro-as-its-flagship-model-leaves-preview\u002F","fead4c82-820e-4f41-a506-bfd417cdacf9",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"b52db7e9-7c58-42c3-9536-5132cb2f8f72","deepseek",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"e4aa9bda-a94d-4299-a0fa-ce5521cf11d2","en","DeepSeek V4 Pro GA: 1.6T MoE with 1M context","DeepSeek V4 Pro 0813 went GA on OpenRouter on August 12, 2026. The 1.6T-total \u002F 49B-active MoE ships with a 1M-token context window, 384K max output, and the preview-era API price of $0.435 per million input tokens and $0.87 per million output tokens. The Pro endpoint is capped at 500 concurrent requests, well below the 2,500 ceiling on Flash.","## Background\n\nDeepSeek opened both V4 Pro and V4 Flash weights under the MIT license and brought them onto its own API on April 24, 2026. From that day the release was framed in two stages: a long public preview, then a separate GA. Flash went official on July 31. Pro kept the preview label. Four months later, the Pro endpoint also goes official, under build number **0813**.\n\nV4 Pro is not a routine dense upgrade. It is a Mixture-of-Experts with 1.6T total parameters and 49B active per token. The architecture wires two new attention variants into the same model: DeepSeek calls them **Compressed Sparse Attention (CSA)** and **Heavily Compressed Attention (HCA)**. The vendor's own ledger says that, at the 1M-token setting, single-token inference compute drops to 27 percent of what V3.2 needed, and KV cache drops to 10 percent. Both V4 models were pre-trained on more than 32T tokens; post-training grew domain-specific experts separately and then consolidated them back into one model through on-policy distillation.\n\nThe preview repo on Hugging Face has logged more than 1.4M downloads in the last month. The model card recommends a context of at least 384K tokens when running the model at maximum reasoning effort locally.\n\n## What ships in 0813\n\n**1) API economics carry over from preview**\n\n- Input: **$0.435 per million tokens** on cache miss, **$0.003625 per million tokens** on cache hit\n- Output: **$0.87 per million tokens**\n- Context window: **1,048,576 tokens** (1M), max output **384,000 tokens**\n- Pro endpoint concurrency cap: **500**; Flash endpoint: **2,500** — Pro remains the slow-but-heavy line\n\nOn OpenRouter, the deepseek-v4-pro-0813 page is stamped \"Released Aug 12, 2026,\" matching the deepseek-v4-pro version string in DeepSeek's own API documentation.\n\n**2) Three operating modes**\n\nDeepSeek splits the API into three effort levels:\n- non-thinking\n- high reasoning effort\n- max effort (the docs describe it as \"pushing the boundary of model reasoning capability\")\n\nThe API is OpenAI ChatCompletions-compatible, Anthropic Messages-compatible, and also exposed through DeepSeek's own Responses API. Both Pro and Flash support tool calling and JSON outputs.\n\n**3) Vendor-reported benchmarks (V4-Pro-Max mode)**\n\n- SWE-bench Verified: **80.6%** resolved\n- Terminal Bench 2.0: **67.9%** accuracy\n- GPQA Diamond: **90.1%** pass@1\n- Humanity's Last Exam: **37.7%** pass@1\n- MMLU-Pro: **87.5%**\n- LiveCodeBench: **93.5%** pass@1\n- Codeforces rating: **3,206**\n- MRCR @ 1M tokens: **83.5 MMR**\n\nThe card's own comparison table places V4-Pro-Max behind GPT-5.4 xHigh on Terminal Bench 2.0 (67.9 vs. 75.1) and behind Gemini-3.1-Pro on Humanity's Last Exam (37.7 vs. 44.4). On SWE-bench Verified (80.6) it lands level with Gemini-3.1-Pro, a hair behind Claude Opus 4.6 (80.8). It takes the top LiveCodeBench and Apex Shortlist scores in the table. None of these numbers has yet been independently replicated for the 0813 build.\n\n**4) Commercial side: Pro pricing will rise**\n\nDeepSeek's pricing page carries a notice that \"a significant increase\" in overall API pricing is coming \"in the near future,\" with specifics to follow in an official announcement. Until that lands, the August 12 list prices hold. On OpenRouter, the average price users actually pay is well below the $0.435 list price, which the platform attributes to caching and discounts.\n\n**5) Open weights for 0813 are not out**\n\nThe Hugging Face repositories still host the April preview builds, and the model card's download table points at those artifacts. DeepSeek has not announced a timeline for publishing 0813 weights, nor confirmed whether the GA build differs from preview beyond post-training. The V4 cadence the company has stated is clear: **API first, weights later.**\n\n## What it means\n\nV4 Pro's GA is not a technical surprise. It is DeepSeek publicly setting a product position.\n\nFirst, **API-first, weights-later** has been DeepSeek's stable release pattern since V3. Stripping the \"preview\" label merely locks down the commercial parameters — price, concurrency, effort levels — while weights trickle out in batches. It is a textbook \"claim the slot first, let the community do the porting and benchmark replication second\" play.\n\nSecond, the **CSA + HCA pairing and the 27% \u002F 10% compute and KV ledger** are what pull \"long context\" back into the MoE camp. In the V3.2 era, KV growth in MoE made anything beyond roughly 100K tokens a non-starter for many teams. V4 Pro uses structural sparsity plus heavy compression to add a \"1M tokens still runs\" secondary track on top of the MoE mainline. Paired with OpenAI and Anthropic pushing GPT-5.6 and the Claude Opus line to comparable 1M context in the same window, the takeaway is: **in the second half of 2026, 1M tokens is an entry ticket for frontier models, not a differentiator.**\n\nThird, the **\"significant price increase\"** notice signals that, with a 500-billion-yuan round reportedly in motion, DeepSeek is shifting from \"win market share on price\" to \"sell API on capability.\" With V4 Flash 0731 already landing within a point of GPT-5.6 Luna on the Artificial Analysis Intelligence Index, the room for Pro to raise prices comes from V4-Pro-Max's positioning on developer workflows — SWE-bench and LiveCodeBench — where it has caught or passed GPT-5.4 and Claude Opus 4.6. Developer workloads tolerate higher per-token prices far better than chat or summarization do.\n\nThe question worth holding open: **as long as the 0813 weights stay locked, DeepSeek is using the wording gap between preview and GA as a commercial buffer. The moment weights and GA land together, the next round of MoE long-context competition gets picked up immediately by Kimi, Qwen, and ByteDance's open-weight flagships.** That window is roughly the next 30 days.\n\n---\n\n**Sources**\n\n- Unite.AI (2026-08-12): *DeepSeek Ships V4 Pro as Its Flagship Model Leaves Preview* — https:\u002F\u002Fwww.unite.ai\u002Fdeepseek-ships-v4-pro-as-its-flagship-model-leaves-preview\u002F\n- OpenRouter model page (2026-08-12): *DeepSeek V4 Pro 0813 — API Pricing & Benchmarks* — https:\u002F\u002Fopenrouter.ai\u002Fdeepseek\u002Fdeepseek-v4-pro-0813\n- Hugging Face (April 2026 onward, continuously updated): *deepseek-ai\u002FDeepSeek-V4-Pro* model card — https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4-Pro","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00Z","2026-08-13T14:05:52.718212Z","2026-08-19T01:48:03.231362Z",true,"agent",104,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"c71b8ee7-9487-4c78-89fd-30bb0368b99e","DeepSeek V4 Flash：284B\u002F13B MoE，成本比 Luna 低 60%","deepseek-v4-flash-0731-intelligence-index-50","2026-08-05T03:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"3cc63477-1334-497d-80cb-90850c019101","DeepSeek-V4-Flash 转正:不靠换架构,只做后训练重新发力 Agent","deepseek-v4-flash-official-post-training-agent-0731","2026-07-31T08:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"80315de0-7eb3-491a-b2e6-103a691a8bd7","Nanbeige4.2-3B 用 Looped Transformer 在 11 项基准上跑赢 Qwen3.5-9B","nanbeige-4-2-3b-looped-transformer-agentic-3b-beats-qwen3-5-9b","2026-07-30T10:30:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00"]