[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepseek-v4-pro-csa-hca-1456-elo-27pct-flops":3,"news-related-0620b8c4-65be-4230-8adb-956c282bdc8b":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"0620b8c4-65be-4230-8adb-956c282bdc8b","DeepSeek V4-Pro 代码能力跃升至第三：压缩注意力机制如何重写百万级上下文效率","4月24日，DeepSeek 发布 V4-Pro（1.6T\u002F49B 激活）与 V4-Flash（284B\u002F13B 激活），核心变化是引入压缩稀疏注意力（CSA）+ 重度压缩注意力（HCA），配合 mHC 超连接，在 100 万 token 上下文下将单 token 推理 FLOPs 压至 V3.2 的 27%，KV Cache 压至 10%。\n\n长上下文推理曾是开源模型禁区——KV 存储随上下文线性增长，128K 是大多数模型的极限。V4 把这道墙凿穿了。两者均支持 Thinking\u002FNon-Thinking 双模式，输出最长 384K，基础上下文窗口统一为 100 万token。\n\nArena AI 代码榜单上，V4-Pro Thinking 以 1456 Elo 排名第三（仅次于 GLM-5.1 的 1534 和 Kimi K2.6 的 1529），Codeforces 评分 3206，超越 GPT-5.4 xHigh 的 3168——开源模型首次在竞争级编程榜单上实质性领先闭源前沿模型。但 MRCR 1M 检索（83.5 vs Opus 4.6 的 92.9）表明长上下文精确检索仍是 Opus 的主场，V4 的优势在于效率而非全面超越。\n\nFlash-Max 性价比尤为突出：输出价格仅 $0.28\u002FM（Pro 为 $3.48\u002FM），LiveCodeBench 91.6 与 Pro 版差距极小。从 MIT 升级到 Apache 2.0 也为企业商业化部署提供了更清晰的专利保护。\n\nDeepSeek V4 最值得关注的不只是跑分，而是效率曲线被彻底改变。27% FLOPs 和 10% KV Cache 意味着超长上下文推理成本第一次可以和短上下文模型相提并论，对需要处理长代码库、长文档的应用是实质性利好。开源社区等了许久的\"百万 token 随便用\"时代，或许就从这里开始。","https:\u002F\u002Fx.com\u002Fdeepseek_ai\u002Fstatus\u002F2047516945466188072","4194681c-1a38-405d-a917-40e1dc2622ea",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"f2bb04db-ef10-4a30-8c0b-e313b27429c1","en","DeepSeek V4-Pro climbs to #3 in coding on compressed attention","On April 24, DeepSeek released V4-Pro (1.6T\u002F49B activated) and V4-Flash (284B\u002F13B activated), with the core change being the introduction of Compressed Sparse Attention (CSA) + Heavy Compressed Attention (HCA), combined with mHC hyper-connections, reducing per-token inference FLOPs to 27% of V3.2 and KV Cache to 10% under 1 million token context.\n\nLong-context inference used to be an open-source model禁区 — KV storage grows linearly with context, 128K is most models' limit. V4 has broken through this wall. Both support Thinking\u002FNon-Thinking dual modes, with max output of 384K, base context window unified at 1 million tokens.\n\nOn the Arena AI coding leaderboard, V4-Pro Thinking ranks third with 1456 Elo (behind GLM-5.1 at 1534 and Kimi K2.6 at 1529), Codeforces score 3206, surpassing GPT-5.4 xHigh's 3168 — the first time an open-source model has substantively led closed-source frontier models on competitive programming leaderboards. But MRCR 1M retrieval (83.5 vs Opus 4.6's 92.9) shows long-context precise retrieval is still Opus's home turf; V4's advantage is efficiency, not comprehensive surpassing.\n\nFlash-Max's cost-performance is particularly outstanding: output price only $0.28\u002FM (Pro is $3.48\u002FM), LiveCodeBench 91.6 with minimal gap from the Pro version. The upgrade from MIT to Apache 2.0 also provides clearer patent protection for enterprise commercial deployment.\n\nWhat's most noteworthy about DeepSeek V4 isn't just the benchmarks, but that the efficiency curve has been completely changed. 27% FLOPs and 10% KV Cache mean ultra-long-context inference cost can for the first time be on par with short-context models — a substantive positive for applications needing to handle long codebases, long documents. The \"million tokens, use at will\" era the open-source community has long awaited may begin here.","deepseek-v4-pro-csa-hca-1456-elo-27pct-flops","2026-04-27T01:00:00Z","2026-04-27T01:08:29.664489Z","2026-08-19T02:08:40.142862Z",true,"agent",104,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7cacddc6-fa02-4de9-84a2-c3320e225571","因果归因剪枝 CAP：让 LLM 推理能力不再随稀疏化而流失","cap-causal-attribution-pruning-arc-61pct","2026-06-20T22:14:08.915874+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"9a1e1c85-60eb-47c6-92b5-bace1746e217","大模型竞争进入下半场：从「比参数」到「比部署」——2026年5月技术格局观察","llm-2nd-half-deploy-vs-params-may-2026","2026-05-25T05:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7dbe12ab-8a86-4e19-a849-b6b0be3f985c","Qwen3.7-Max评测揭示推理代价：97M token输出背后的效率博弈","qwen3-7-max-97m-tokens-extended-thinking","2026-05-22T10:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"9bb023ae-147a-4081-a973-5638e260803f","1M 上下文实测：Gemini 3.1 Pro 与 Opus 4.7 稳，GPT-5.5 在 512K 衰减","1m-context-multihop-benchmark-cliff-degradation","2026-05-15T22:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"72a30e44-f38d-42af-af4a-32d265f76608","EfficientLLM：大模型效率研究的首次系统性「全景扫描」","efficient-llm-benchmark-panorama-tradeoff","2026-05-14T08:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"e2a935d5-4893-4acb-bdb5-1783c19eeb20","xAI悄然发布Grok 4.3：速度致胜，但智能仍未登顶","grok-4-3-xai-207-tps-cheap-fast","2026-05-03T16:01:00+00:00"]