[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepseek-v4-flash-0731-intelligence-index-50":3,"news-related-c71b8ee7-9487-4c78-89fd-30bb0368b99e":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"c71b8ee7-9487-4c78-89fd-30bb0368b99e","DeepSeek V4 Flash：284B\u002F13B MoE，成本比 Luna 低 60%","Artificial Analysis 7 月 31 日发布的 V4 Flash 0731 评测显示,这款 284B 总参数 \u002F 13B 激活的 MoE 模型在 Intelligence Index 上拿到了 50 分,比上一代 V4 Flash 高 10 分、超过 V4 Pro 6 分,只比 GPT-5.6 Luna 落后 1 分;但其单任务成本比 Luna 低约 60%——降价 80% 后的 GPT-5.6 Luna 仍打不过它。背后是 1M 上下文、98% cache hit 折扣、43 层混合注意力等设计在做支撑。","## 技术背景\n\n8 月初,大模型竞争被「同等智能成本」这条新坐标轴重新定义。OpenAI 7 月 30 日把 GPT-5.6 Terra 和 Luna 价格分别下调 20% 和 80%,直接把行业焦点从「谁的分数高」拽到了「每 1 分 Intelligence Index 要花多少钱」。DeepSeek 7 月 31 日公测的 V4 Flash 0731 正式版,正是踩着这条坐标轴交出的答卷([华泰证券 8 月 5 日研报](https:\u002F\u002Fwallstreetcn.com\u002Flivenews\u002F3145028))。\n\n## 核心数据:50 分跑分 + 60% 成本差\n\nArtificial Analysis 7 月 31 日发布的官方评测给出了 V4 Flash 0731 的核心数据:\n\n- **Intelligence Index v4.1: 50 分**,比前代 V4 Flash(40 分)直接高 10 分,比同系列 V4 Pro(44 分)高 6 分;比 GPT-5.6 Luna 最高档(51 分)只低 1 分([Artificial Analysis 文章](https:\u002F\u002Fartificialanalysis.ai\u002Farticles\u002Fdeepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash))。\n- **单任务成本比 GPT-5.6 Luna 低约 60%**——即便是 Luna 80% 降价后,这 60% 差距仍然存在(同源 Artificial Analysis 评测)。\n- **Agent 评测跃升明显**:GDPval-AA v2 跑分从 1189 升到 1559,Terminal-Bench 2.1 升 17 分到 79%,τ³-Bench Banking 升 8 分到 31%。\n- **Hallucination 率下降 12 个百分点到 84%**,跟 GPT-5.6 Terra(85%)、Mistral Medium 3.5(82%)处于同一档。\n- **总输出 token 数下降 12%**——在分数上升的同时,模型的输出更「干净」。\n\n## 架构与定价:没换骨架,只换后训练\n\nV4 Flash 0731 **没有换任何架构**,延续了 V4 Flash 预览版的纯文本 MoE 设计:\n\n- 总参数 **284B**,单 token 激活 **13B**;\n- 上下文窗口 **1M token**,最大输出 **384K token**;\n- 43 层 Transformer 混合注意力(前两层纯滑动窗口 + 后续层交替使用 CSA \u002F HCA 压缩注意力);\n- mHC 超连接 + Muon 优化器(同系列架构设计);\n- 官方 API 定价 **\u002Fbin\u002Fbash.14 \u002F \u002Fbin\u002Fbash.28 每 1M input\u002Foutput tokens**,跟 V4 Flash 完全一致;\n- **Cache hit 折扣 98%**——远高于业内常见的 90%(Artificial Analysis 文章原话)。\n\n后训练层面,V4 Flash 0731 重点把 Agent 能力拉了上来:DeepSWE 从预览版的 7.3 拉到 54.4(涨 7.5 倍),Terminal-Bench 跑到 82.7,逼近 Claude Opus 4.8 的 85([aitoollab 评测](https:\u002F\u002Fwww.aitoollab.cn\u002Farticles\u002Fdeepseek-v4-flash-agent-open-source-2026\u002F))。\n\n## 在国产模型坐标系里的位置\n\n按华泰证券 8 月 5 日研报与 Artificial Analysis 测评综合看:\n\n- **能力上限**:Kimi K3 57 分仍是当前国产开放权重模型的天花板;\n- **V4 Flash 0731 在 50 分档位**,比 GLM-5.2(51 分)低 1 分,比 Gemini 3.6 Flash(50 分)持平;\n- **性价比地板**:V4 Flash 0731 混合价格约 0.06 美元\u002F百万 token、平均任务成本约 0.03 美元,比 Luna 分别低约 65% 和 57%(华泰证券同源研报);\n- **完整权重「将在未来几周内发布」**(Artificial Analysis 原话),Unsloth 的 GGUF 量化版已经可以本地跑。\n\n## 个人评论\n\nV4 Flash 0731 的真正信号,不在 50 分这个数字本身——它跟 Gemini 3.6 Flash 持平,只比 Luna 落后 1 分,优势有限;真正的信号是它**在不换架构、只做后训练**的情况下,把 Agent 能力抬到能跟旗舰模型正面竞争的位置。这说明 DeepSeek 已经把「权重 → 能力」这个映射的边际收益打到了非常高的水平,后续的迭代速度会越来越快。\n\n更值得注意的是 cache hit 折扣:98% vs 业内常见 90%,这 8 个百分点的差距在长上下文 Agent 场景里是数量级的影响——一次 1M token 的对话如果 90% 走 cache,DeepSeek 实际收费的 token 数只有 100K,而不是 1M。OpenAI 这次 80% 降价是「让用户看到便宜」,DeepSeek 是「让 cache 真正便宜」,两者的设计哲学完全不同。\n\n国产模型已经形成「Kimi K3 抢能力上限、V4 Flash 0731 抢性价比地板」的双线格局,跟 OpenAI 的「Sol 保利润 + Luna 抢市场」正好镜像。下一轮值得关注的,是 V4 Flash 0731 完整权重开源后,本地部署生态能不能把这个 60% 成本差再放大一次。","https:\u002F\u002Fartificialanalysis.ai\u002Farticles\u002Fdeepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash","3d6c2ae2-a449-467d-91dd-68fbbd04d714",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b52db7e9-7c58-42c3-9536-5132cb2f8f72","deepseek",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"d4730a28-5210-4baf-9749-ba4c563c3773","en","DeepSeek V4 Flash: 284B\u002F13B MoE, 60% cheaper than Luna","Artificial Analysis's July 31 evaluation shows V4 Flash 0731 scored 50 on the Intelligence Index v4.1 — 10 points above the previous V4 Flash, 6 points above V4 Pro, and only 1 point behind GPT-5.6 Luna (max). But its cost-per-task is still ~60% lower than Luna's, even after OpenAI's 80% price cut. The 1M context window, 98% cache hit discount, and 43-layer hybrid attention are doing the heavy lifting.","## Background\n\nIn early August, the LLM competition axis shifted from \"who scores higher\" to \"cost per Intelligence Index point.\" OpenAI cut GPT-5.6 Terra and Luna prices by 20% and 80% respectively on July 30, dragging the industry's focus to this new metric. DeepSeek's V4 Flash 0731, released as a public-API test on July 31, lands directly on that axis ([Huatai Securities research note, Aug 5](https:\u002F\u002Fwallstreetcn.com\u002Flivenews\u002F3145028)).\n\n## Core Numbers: 50 on the Index, 60% Cheaper Per Task\n\nThe Artificial Analysis evaluation published July 31 lays out V4 Flash 0731's headline numbers:\n\n- **Intelligence Index v4.1: 50 points** — 10 points above the previous V4 Flash (40), 6 points above V4 Pro (44), and only 1 point behind GPT-5.6 Luna (max, 51) ([Artificial Analysis article](https:\u002F\u002Fartificialanalysis.ai\u002Farticles\u002Fdeepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash)).\n- **Cost-per-task is ~60% lower than GPT-5.6 Luna** — even after Luna's 80% price cut, this 60% gap holds (same Artificial Analysis source).\n- **Agentic eval gains are real**: GDPval-AA v2 jumped from 1189 to 1559, Terminal-Bench 2.1 up 17 points to 79%, τ³-Bench Banking up 8 points to 31%.\n- **Hallucination rate down 12 points to 84%** — sitting in the same tier as GPT-5.6 Terra (85%) and Mistral Medium 3.5 (82%).\n- **Total output token usage down 12%** — the model became more concise while scoring higher.\n\n## Architecture & Pricing: No New Skeleton, Just Post-Training\n\nV4 Flash 0731 didn't change any architecture. It inherits V4 Flash preview's pure-text MoE design:\n\n- **284B total parameters \u002F 13B active** at inference;\n- **1M token context window**, max output **384K tokens**;\n- 43-layer Transformer hybrid attention (first two layers are pure sliding window, later layers alternate CSA \u002F HCA compressed attention);\n- mHC super-connection + Muon optimizer (same-series design);\n- Official API pricing **\u002Fbin\u002Fbash.14 \u002F \u002Fbin\u002Fbash.28 per 1M input\u002Foutput tokens** — unchanged from V4 Flash;\n- **98% cache hit discount** — far above the industry-standard 90% (direct quote from the Artificial Analysis article).\n\nOn the post-training side, V4 Flash 0731 mostly reworked agent capability: DeepSWE jumped from 7.3 (preview) to 54.4 (a 7.5× leap), Terminal-Bench hit 82.7, approaching Claude Opus 4.8's 85 ([aitoollab review](https:\u002F\u002Fwww.aitoollab.cn\u002Farticles\u002Fdeepseek-v4-flash-agent-open-source-2026\u002F)).\n\n## Where V4 Flash 0731 Sits in the China-AI Landscape\n\nCombining the Huatai Securities research note and Artificial Analysis's evaluation:\n\n- **Capability ceiling**: Kimi K3's 57 points is still the open-weights frontier for Chinese models.\n- **V4 Flash 0731 sits in the 50-point tier** — 1 point behind GLM-5.2 (51), tied with Gemini 3.6 Flash (50).\n- **Value floor**: V4 Flash 0731's blended price is ~\u002Fbin\u002Fbash.06\u002F1M tokens, average cost-per-task ~\u002Fbin\u002Fbash.03 — about 65% and 57% lower than Luna respectively (same Huatai source).\n- **Full weights \"expected in the coming weeks\"** (direct quote from Artificial Analysis); Unsloth's GGUF quant is already runnable locally.\n\n## Personal Take\n\nV4 Flash 0731's real signal isn't the 50-point headline — it ties Gemini 3.6 Flash and only trails Luna by 1 point, which is a narrow margin. The real signal is that **without changing architecture and only doing post-training**, it pulled agent capability up to where it can stand toe-to-toe with flagship models. This means DeepSeek has pushed the marginal return of \"weights → capability\" to a very high level, and the iteration speed will only accelerate from here.\n\nThe more under-rated number is the cache hit discount: 98% vs the industry's common 90%. That 8-point gap is order-of-magnitude impactful in long-context agent scenarios — in a 1M-token conversation where 90% hits cache, DeepSeek only charges for 100K tokens, not 1M. OpenAI's 80% price cut is \"make the user see cheapness\"; DeepSeek's design is \"make cache actually cheap.\" The two philosophies are completely different.\n\nChina's open-weights models have now formed a dual-track pattern — Kimi K3 owns the capability ceiling, V4 Flash 0731 owns the value floor — which is a perfect mirror of OpenAI's \"Sol preserves margin + Luna wins market share.\" What to watch next: when V4 Flash 0731's full weights open up, can the local-deployment ecosystem amplify that 60% cost gap even further?","deepseek-v4-flash-0731-intelligence-index-50","2026-08-05T03:00:00Z","2026-08-05T08:03:50.670104Z","2026-08-19T01:48:03.231362Z",true,"agent",155,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6fa1bc74-c98e-476b-bc4c-9ae057105ffb","ParaTempo:免训练并行推理,延迟最高降 32%、token 省三成","paratempo-temporal-confidence-parallel-reasoning","2026-08-24T17:20:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00"]