[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepseek-v4-flash-vision-exp-multimodal":3,"news-related-12a5f49d-8c83-4c40-82af-1c0b7f1c8b3e":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"12a5f49d-8c83-4c40-82af-1c0b7f1c8b3e","DeepSeek 给 V4-Flash 装上眼睛:Vision-Exp 实验模型两项基准反超 Opus 4.8","DeepSeek 发布实验性多模态模型 V4-Flash-Vision-Exp,在 V4-Flash 文本能力上叠加图像理解,官方基准显示其在 Agents' Last Exam、ZeroBench 两项多模态测试反超 Claude Opus 4.8。","在 V4-Pro 与 V4-Flash 双旗舰跑了四个月之后,DeepSeek 给自家「小快灵」那条产品线补上了眼睛。8 月 21 日,DeepSeek 上线实验性多模态模型 V4-Flash-Vision-Exp:它在 DeepSeek-V4-Flash 的文本能力之上叠加图像理解。官方定位不是新旗舰,而是面向 agent 应用——把视觉理解与工具使用结合起来,给需要「边看边干」的工作流用。\n\n## 先看成绩单:两项反超,一项还在追\n\n据 DeepSeek 官方公布的对比数据(经 OfficeChai 整理),新模型在文本类 agent 评测上与上一版 V4-Flash-0731 几乎持平,这印证了官方「视觉版不牺牲文本能力」的说法。Terminal Bench 2.1 上,V4-Flash-Vision-Exp 拿到 83.9,略高于旧版 V4-Flash 的 82.7,距离 Opus 4.8 的 85.0 只差 1.1 分;但 NL2Repo 上差距被拉大到 57.7 对 69.7,DSBench-Hard 也落后约 8 分。\n\n真正有信息量的是多模态侧:ApexBench(Pass@1)上视觉版拿到 36.5,对比文本版被迫忽略图像输入时的 26.2,提升了 10 分以上——虽然仍落后 Opus 4.8 的 39.4。而在 Agents' Last Exam(27.3 对 25.7)和 ZeroBench Pass@5(35.0 对 34.0)两项上,V4-Flash-Vision-Exp 直接反超了 Anthropic 的现役旗舰。Chartography 上 64.3 对 65.0,基本咬平。\n\n需要提醒的是,这些数字出自 DeepSeek 自家的 Harness Minimal Mode 测试环境(top_p 0.95、temperature 1.0),并非独立复测,参考时请保留一分克制。\n\n## 工程细节比跑分更值得关注\n\n这次同步放出的还有一批面向开发者的基建([官方 Vision 指南](https:\u002F\u002Fapi-docs.deepseek.com\u002Fguides\u002Fvision\u002F)):\n\n- **按内容识别格式**:模型支持 JPEG、PNG、GIF、WebP,格式判定看文件实际字节而非文件名或 MIME 声明;\n- **384 token 封顶计费**:每张图最多计 384 token,定价沿用 V4-Flash 费率;可选 detail 字段把图降采样到 512×512 省 token,处理前自动归一化到约 800×800;\n- **Files API 免费**:上传一次图像、按 ID 跨请求复用(单文件 64 MiB),不必每次重传;单请求最多 600 张图,单边最长 8192 像素;\n- **三协议兼容**:同时支持 OpenAI 的 Chat Completions\u002FResponses 和 Anthropic 的 Messages 端点,同日发布的 Harness 0.1.1 开箱支持新模型。\n\n## 实验模型是一步便宜的棋\n\nOfficeChai 的分析点破了这次发布的策略属性:DeepSeek 没有从零训练新旗舰,而是给已经便宜的 V4-Flash 加视觉,用最小成本把产品线延伸进多模态 agent 领域。同样的打法在 MiniMax、智谱身上也能看到:多模态能力正在下沉到中端模型,而不再是大参数旗舰的专属卖点。\n\n对开发者的启示很直接:当「能看图的便宜模型」开始逼近旗舰的多模态基准,选型逻辑就要从「谁跑分最高」切换到「单位成本下的视觉 agent 能力」。实验标签意味着它可能随时调整或毕业,适合先在低风险工作流里试水,而不是直接押上生产关键路径。\n\n下一轮中国实验室的竞争,未必是谁的模型最大,而是谁能让便宜快速的模型先学会看世界——这一局,DeepSeek 先出了牌。","https:\u002F\u002Fapi-docs.deepseek.com\u002Fguides\u002Fvision\u002F","4194681c-1a38-405d-a917-40e1dc2622ea",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"b52db7e9-7c58-42c3-9536-5132cb2f8f72","deepseek",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"d969770d-53b7-4782-8ded-1669d3084979","en","DeepSeek Gives V4-Flash Eyes: Vision-Exp Beats Opus 4.8 on Two Benchmarks","DeepSeek has released V4-Flash-Vision-Exp, an experimental multimodal model that adds image understanding to V4-Flash's text capabilities. On the company's own benchmarks it edges past Claude Opus 4.8 on Agents' Last Exam and ZeroBench.","Four months after shipping its V4-Pro and V4-Flash flagship pair, DeepSeek has given its budget-speed line a pair of eyes. On August 21, DeepSeek launched V4-Flash-Vision-Exp, an experimental multimodal model that layers image understanding on top of DeepSeek-V4-Flash's text capabilities. It is positioned not as a new flagship but as infrastructure for agent applications — combining visual understanding with tool use for workflows that need to see while they act.\n\n## The Scorecard: Two Wins, One Still Chasing\n\nAccording to DeepSeek's own comparison data (compiled by OfficeChai), the new model lands close to the previous V4-Flash-0731 on text-based agent evaluations, supporting the company's claim that the vision variant doesn't sacrifice text performance. On Terminal Bench 2.1, V4-Flash-Vision-Exp scores 83.9 versus 82.7 for the older V4-Flash, just 1.1 points behind Opus 4.8's 85.0. But the gap widens on NL2Repo — 57.7 against 69.7 — and DSBench-Hard trails Opus 4.8 by roughly eight points.\n\nThe multimodal side is where the release earns its billing. On ApexBench (Pass@1), the vision model scores 36.5, a jump of more than ten points over the 26.2 the text-only V4-Flash manages when forced to ignore image inputs — though still behind Opus 4.8's 39.4. On Agents' Last Exam (27.3 vs 25.7) and ZeroBench Pass@5 (35.0 vs 34.0), V4-Flash-Vision-Exp actually edges past Anthropic's current flagship. On Chartography, 64.3 vs 65.0 is nearly a tie.\n\nOne caveat: these numbers come from DeepSeek's own Harness Minimal Mode setup (top_p 0.95, temperature 1.0), not independent verification — treat them with appropriate caution.\n\n## Engineering Details Matter More Than Scores\n\nThe release also ships developer-facing infrastructure ([official Vision guide](https:\u002F\u002Fapi-docs.deepseek.com\u002Fguides\u002Fvision\u002F)):\n\n- **Content-based format detection**: the model handles JPEG, PNG, GIF and WebP, determining the format from actual file bytes rather than filename or declared MIME type;\n- **384-token cap per image**: each image costs at most 384 tokens at V4-Flash's existing rates; an optional detail field downscales images to 512 x 512 to save tokens, with automatic normalization to roughly 800 x 800 before processing;\n- **Free Files API**: upload an image once and reference it by ID across requests (64 MiB per file), with up to 600 images per request and 8,192-pixel max edge length;\n- **Triple-protocol support**: OpenAI's Chat Completions and Responses APIs plus Anthropic's Messages endpoint all work, and Harness 0.1.1 released the same day supports the new model out of the box.\n\n## An Experimental Model Is a Cheap Move\n\nOfficeChai's analysis cuts to the strategic core: rather than training a new flagship from scratch, DeepSeek added vision to the already-cheap V4-Flash, extending its lineup into multimodal agent territory at minimal cost. The same playbook is visible at MiniMax and Zhipu AI — multimodal capability is sinking into mid-tier models instead of remaining an exclusive selling point for giant flagships.\n\nThe takeaway for developers is direct: when a cheap model that can see approaches flagship-level multimodal benchmarks, model selection shifts from \"who scores highest\" to \"visual agent capability per unit cost.\" The experimental label means it could be adjusted or graduate at any time — sensible for low-risk workflow trials, not yet for betting the production critical path.\n\nThe next round of competition among Chinese AI labs may not be about who has the biggest model, but who can teach a cheap, fast one to see the world first. DeepSeek has played its card.","deepseek-v4-flash-vision-exp-multimodal","2026-08-21T23:05:00Z","2026-08-21T23:07:08.708991Z","2026-08-21T23:07:08.708999Z",true,"agent",76,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"3cc63477-1334-497d-80cb-90850c019101","DeepSeek-V4-Flash 转正:不靠换架构,只做后训练重新发力 Agent","deepseek-v4-flash-official-post-training-agent-0731","2026-07-31T08:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"7258978b-dfcd-4cb4-91c4-3b8569cd5deb","Qwen-Audio-3.0-TTS双版本发布:Plus登顶Artificial Analysis,Flash压到300ms首包延时","qwen-audio-3-tts","2026-07-20T10:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"652b1484-31eb-4a83-a932-21fcf97b3a50","Boson AI Higgs Audio v3 TTS：4B 参数原生可控百语种语音生成","higgs-audio-v3-boson-4b-100-language","2026-06-04T18:00:00+00:00"]