[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-vision-exp-benchmark-reread-3-wins":3,"topics-all":38,"news-related-47bdcf73-18de-439d-ab05-a0666533d360":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"47bdcf73-18de-439d-ab05-a0666533d360","Vision-Exp权重基准重读:3项反超Opus,Chartography差0.7分","8月31日DeepSeek开源Vision-Exp权重(305B、MIT),其11项基准表逐项重读:反超Opus-4.8的是DeepSWE、ZeroBench、Agents' Last Exam共3项,Chartography以64.3对65.0居次,附vLLM\u002FSGLang部署配方。","8月31日，DeepSeek 把 V4 家族第一个多模态实验模型 DeepSeek-V4-Flash-Vision-Exp 的权重放上了 Hugging Face——距离它以 API 形式上线（8月21日）只过了10天。305B 参数、MIT 许可证，附极简 PyTorch 推理参考实现。权重开放之后最值得做的事，是把官方那张 11 项基准对照表逐项重读一遍：哪些是真反超，哪些其实没有。\n\n## 3 项反超 Opus，但不在 Chartography\n\n模型卡的三方对照（Vision-Exp \u002F 前代 0731 文本版 \u002F Anthropic Opus-4.8）覆盖 7 项文本 agent 基准与 4 项多模态基准。逐项核对，Vision-Exp 真正压过 Opus-4.8 的只有 3 项：文本侧 DeepSWE 59.3 对 58.0；多模态侧 ZeroBench（Pass@5）35.0 对 34.0、Agents' Last Exam 27.3 对 25.7。\n\n而常被引用的 Chartography 实际是 64.3 对 65.0——Opus 以 0.7 分领先；ApexBench（Pass@1）36.5 对 39.4 同样 Opus 领先。文本侧 Terminal Bench 2.1（83.9 对 85.0）、NL2Repo（57.7 对 69.7）、Cybergym（75.3 对 78.3）、DSBench-Hard（63.6 对 71.7）也都是 Opus 占优。差距是结构性的：纯文本重任务 Opus 仍领先，视觉相关与 SWE 场景 Vision-Exp 开始咬住甚至反超。\n\n与前代 0731 相比，文本侧 7 项里 6 项提升：DeepSWE 从 54.4 到 59.3、Toolathlon-Verified 从 70.3 到 75.9、DSBench-Hard 从 59.6 到 63.6、Terminal Bench 2.1 从 82.7 到 83.9，唯一微降的是 Cybergym。多模态侧跃升更猛：ApexBench 从 26.2 到 36.5——表格脚注注明 0731 在这项忽略了输入中的多模态元素，这 10.3 分正是「装上眼睛」的直接收益。\n\n## 权重仓库里还有什么\n\n参考实现覆盖视觉编码器、aligner、DFlash 注意力、MoE、Hyper-Connections 和 DSpark 前向路径——组件名对应 DeepSeek 自家今年的注意力工作，非现成模块缝合。评测配置透明：DeepSeek Harness minimal 模式、max 推理强度、temperature = 1.0、top_p = 0.95。\n\n部署给到命令行级：vLLM 一条 docker 命令可在单台 4×GB300 节点起服务，默认开 DSpark 投机解码（3 个投机 token）；SGLang 侧草稿与目标权重同源一份 checkpoint，无需单独 draft 模型。生态跟进快：上线约两天月下载 17,893，已有 2 个微调、8 个量化版本，llama.cpp\u002FLM Studio\u002FOllama 入口就位。\n\n## 对开发者的所以呢\n\nAPI 阶段的跑分只能听官方一面之词；权重落地后，任何人都能在相同配置下重跑这张表——包括亲手验证到底哪几项反超。MIT 许可证没有商用顾虑。下一个信号：Exp 后缀去掉之时，大概率就是 V4 家族多模态正式版的发布日。\n\n参考：huggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4-Flash-Vision-Exp 模型卡；theopenweights.com\u002Fnews\u002Fdeepseek-v4-flash-vision-exp-2l4q","https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4-Flash-Vision-Exp\u002Fblob\u002Fmain\u002FREADME.md","4194681c-1a38-405d-a917-40e1dc2622ea",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"b52db7e9-7c58-42c3-9536-5132cb2f8f72","deepseek",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"9c22aa3c-49d7-4bbe-8da6-f1ac2cdb2ca0","en","Vision-Exp reread: 3 wins over Opus, Chartography 0.7 short","Item-by-item reread of Vision-Exp's open weights: 3 benchmark wins over Opus-4.8, Chartography 64.3 vs 65.0, plus vLLM\u002FSGLang deployment recipes.","On August 31, DeepSeek put the weights of DeepSeek-V4-Flash-Vision-Exp — the first experimental multimodal model in the V4 family — on Hugging Face, just 10 days after its API launch on August 21. 305B parameters, MIT license, and a minimal PyTorch inference reference implementation. Now that the weights are open, the most useful thing to do is reread the official 11-benchmark table item by item: which wins are real, and which are not.\n\n## Three wins over Opus — but not Chartography\n\nThe model card's three-way comparison (Vision-Exp vs. the previous 0731 text-only release vs. Anthropic Opus-4.8) covers 7 text-agent benchmarks and 4 multimodal ones. Checked line by line, Vision-Exp actually beats Opus-4.8 on exactly three: DeepSWE 59.3 vs 58.0 on the text side; ZeroBench (Pass@5) 35.0 vs 34.0 and Agents' Last Exam 27.3 vs 25.7 on the multimodal side.\n\nThe often-cited Chartography, however, is 64.3 vs 65.0 — Opus ahead by 0.7. ApexBench (Pass@1) is 36.5 vs 39.4, also Opus. On the text side, Terminal Bench 2.1 (83.9 vs 85.0), NL2Repo (57.7 vs 69.7), Cybergym (75.3 vs 78.3), DSBench-Hard (63.6 vs 71.7) and AutomationBench (25.7 vs 27.2) all favor Opus too. The pattern is structural: Opus still leads on heavy text tasks, while Vision-Exp closes in — and occasionally overtakes — on vision-related and SWE ground.\n\nAgainst the 0731 predecessor, six of seven text benchmarks improve: Terminal Bench 2.1 from 82.7 to 83.9, DeepSWE from 54.4 to 59.3, Toolathlon-Verified from 70.3 to 75.9, DSBench-Hard from 59.6 to 63.6, NL2Repo from 54.2 to 57.7, AutomationBench from 25.1 to 25.7; the single dip is Cybergym (76.7 to 75.3). The multimodal jump is larger: ApexBench from 26.2 to 36.5 (+10.3) — note the table footnote: 0731 ignored the multimodal elements in its input for this benchmark, so those 10 points are the direct payoff of adding eyes. Agents' Last Exam goes from 25.2 to 27.3.\n\n## What else is in the repository\n\nThe reference implementation covers the vision encoder, aligner, DFlash attention, MoE, Hyper-Connections, and the DSpark forward path — component names that map to DeepSeek's own stream of attention work this year, not a bolt-on of off-the-shelf modules. Evaluation configuration is explicit: text-agent benchmarks use the minimal mode of DeepSeek Harness with the max reasoning effort level, temperature = 1.0, top_p = 0.95.\n\nDeployment recipes go down to the command line: a single vLLM docker command serves the model on one 4×GB300 node with fp8 KV cache, block size 256, tensor parallel 4, and DSpark speculative decoding on by default (num_speculative_tokens = 3). On SGLang, --speculative-algorithm DSPARK draws draft and target weights from the same checkpoint, so no separate draft model is needed. The ecosystem moved fast too: about two days in, Hugging Face shows 17,893 monthly downloads, 2 finetunes and 8 quantized variants, with llama.cpp, LM Studio and Ollama adapter entries already in place.\n\n## So what for developers\n\nWhile the model is API-only, you have to take the vendor's scores on faith; with weights on disk, anyone can re-run the table under the same configuration — including verifying which benchmarks were actually won. The MIT license carries no commercial strings. The next signal worth watching: when the Exp suffix drops, that is most likely the ship date of the official V4 multimodal release.\n\nReference: huggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4-Flash-Vision-Exp (model card and benchmark table); theopenweights.com\u002Fnews\u002Fdeepseek-v4-flash-vision-exp-2l4q","vision-exp-benchmark-reread-3-wins","2026-09-01T21:08:09Z","2026-09-01T21:08:22.576832Z","2026-09-01T21:08:22.576849Z",true,"agent",142,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"51c13e24-8072-404c-a8d4-75c40cff05ee","Ling-3.0-flash-VL 开源：124B MoE 只激活 5.5B，视觉塞进 Agent 闭环","ling-3-0-flash-vl-open-weights","2026-09-15T13:18:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"550cee5e-18e8-4236-9304-7207ebc221a8","Agnes 3.0 Flash 开源:72 层仅 18 层带 KV 缓存,33B 单卡跑 262k 上下文","agnes-3-0-flash-preview-open-weights","2026-09-13T15:20:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"108af093-b226-4372-9cf0-77323ffc5456","小鹏 X-AuT 给语音大模型剪枝:音频塔砍 4 层,车载推理提速 21.4%","xpeng-x-aut-audio-encoder-pruning","2026-09-12T19:06:47+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"17006864-46a5-405c-a8cc-24507bbc5e37","YuE2-3B 开源:乐谱可编辑的音乐生成,官方基准反超 Suno v5","yue2-3b-editable-music-generation","2026-09-10T13:20:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"cd49f913-cde7-4cf3-8d93-24508653180e","腾讯混元开源AuK:1.5B语音模型统一生成与编辑,4步推理快4.5倍","tencent-hunyuan-auk-speech-editing","2026-09-09T09:12:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"dc2f4ead-963c-4a8e-bd41-400bebf83bb4","物理、几何、外观一个模型全包:Puffin-World 开源,相机 roll 误差低至 0.26°","puffin-world-native-3d-world-states","2026-09-06T19:09:41+00:00"]