[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ling-3-0-flash-vl-open-weights":3,"topics-all":38,"news-related-51c13e24-8072-404c-a8d4-75c40cff05ee":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"51c13e24-8072-404c-a8d4-75c40cff05ee","Ling-3.0-flash-VL 开源：124B MoE 只激活 5.5B，视觉塞进 Agent 闭环","蚂蚁 inclusionAI 开源 Ling 系列首个原生多模态模型：124B 总参数、5.5B 激活、256K 上下文，支持图像视频输入。官方引用 AA 指数 42 分，独立评测为 25 分且站上效率帕累托前沿，但 Terminal-Bench v4.0 得 0 分。","9 月 9 日，蚂蚁集团 inclusionAI 把 Ling 系列第一个原生多模态模型 Ling-3.0-flash-VL 的权重放上了 Hugging Face 和 ModelScope：总参数 124B、每个 token 只激活 5.5B，支持图像和视频输入，上下文窗口 256K，BF16 与 FP8 权重同步开放，FP4 和 INT4 版本按计划后续跟进。这是 Ling 系列从纯文本走向「能看、能想、能动手」闭环的第一步。\n\n## 架构：稀疏 MoE 上面叠一整套混合注意力\n\n模型卡给出的结构并不复杂：ViT 视觉编码器负责抽取图像和视频特征，两层 MLP 投影器把视觉特征对齐到文本表示；VideoRoPE 同时编码空间位置和时间顺序，支撑事件定位、长视频问答和视频片段剪辑这类任务；语言主干是 42 层的 KDA 与 Gated MLA 按 5:1 交替的混合架构，负责长上下文的高效处理；最外层用稀疏 MoE 把总容量抬到 124B，同时把单 token 计算压在 5.5B 激活。换句话说，视觉不是外挂的「看图插件」，而是被塞进了理解、推理、行动、验证的完整流程。\n\n## 跑分：官方引用和独立评测对不上\n\n官方模型卡写的是：Artificial Analysis Intelligence Index v4.1.1 得分 42，比纯文本的 Ling-3.0-flash 的 38 分高 4 分——蚂蚁以此论证「多模态联合训练反过来也提升了文本能力」。但 AA 自己的独立评测帖给出的数字是 25 分，同时把它放在「智能 vs 激活参数」帕累托前沿上：同等总规模里，Qwen3.5 122B A10B 只有 16 分，Mistral Medium 3.5 是 15 分。两个数字大概率对应不同版本的指数，但这也提醒所有人：发布方引用的跑分，永远先看版本号和口径再引用。\n\nAA 的独立评测还给出了官方宣传页不会突出的短板：幻觉率 22%（对比 Inkling Small 的 63% 算克制），但知识召回只有 14% 的 AA-Omniscience 准确率；自动化办公基准 AutomationBench-AA 只有 16%，更难的 Terminal-Bench v4.0 直接 0 分。看图干活的故事讲得再顺，离「可靠的数字员工」还有明显距离。\n\n## 部署现实：5.5B 激活不等于 5.5B 显存\n\n一个容易误读的数字：激活参数 5.5B 压的是单 token 计算量，不是显存占用，部署时仍然要装下 124B 总权重。官方给的参考配置是 4 张 141GB 级显卡（H20-3e\u002FH200 或 B300\u002FGB300 节点）跑 256K 上下文（YaRN 扩展），80GB 的 H100\u002FH800 需要扩到 8 卡张量并行。推理路径两条：SGLang 官方 Docker 镜像，或 inclusionAI 的 vllm-ling-v3 分支走 vLLM。\n\n## 所以呢\n\n开源多模态的竞争已经从「能看图」进入「能干活」的闭环竞争——GUI 操作、前端还原、医疗报告解读是发布方点名的三类场景。5.5B 激活的推理经济学，瞄准的正是闭源厂商走量最大的 flash 层定价。但对使用者来说，真正值得记住的不是官方引用的 42 分，而是独立评测里那个 0 分的 Terminal-Bench：多模态 Agent 的水，还深得很。\n\n参考：[Ling-3.0-flash-VL 模型卡](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-3.0-flash-VL) · [Artificial Analysis 独立评测](https:\u002F\u002Fx.com\u002FArtificialAnlys\u002Fstatus\u002F2098502360800850280) · [Pandaily 报道](https:\u002F\u002Fx.com\u002FthePandaily\u002Fstatus\u002F2097921325578719673)\n","https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-3.0-flash-VL","c90d26bf-0877-40ec-a61a-4697f288f278",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"471c51be-e620-49df-bd6c-0b5504f53f00","ant-group",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"f7484094-01a5-418d-8fb9-1328bb285258","en","Ling-3.0-flash-VL: 124B MoE, 5.5B active, vision in the agent loop","Ant open-sources Ling-3.0-flash-VL: 124B total, 5.5B active, 256K context, image+video input. Official AA index 42; independent eval 25, 0% on Terminal-Bench.","On September 9, Ant Group's inclusionAI open-sourced Ling-3.0-flash-VL, the first natively multimodal model in its Ling line, on Hugging Face and ModelScope: 124B total parameters with only 5.5B activated per token, image and video input, a 256K-token context window, and BF16 plus FP8 weights available immediately, with FP4 and INT4 builds planned to follow.\n\n## Architecture: sparse MoE under a hybrid-attention backbone\n\nThe model card describes a straightforward stack: a ViT visual encoder extracts image and video features, a two-layer MLP projector aligns them with text representations, and VideoRoPE encodes both spatial positions and temporal order — supporting event localization, long-video question answering, and video clip editing. The language trunk is a 42-layer hybrid backbone alternating KDA and Gated MLA layers at a 5:1 ratio for efficient long-context processing, wrapped in a sparse MoE that holds 124B of total capacity while activating 5.5B per token. Vision here is not a bolted-on captioning plugin; the card frames it as integrated into the full loop of understanding, reasoning, acting, and verification.\n\n## Benchmarks: the official citation and the independent eval disagree\n\nThe official card claims 42 on the Artificial Analysis Intelligence Index v4.1.1, four points above the text-only Ling-3.0-flash at 38 — Ant's argument being that joint multimodal training improved text ability as well. But Artificial Analysis's own evaluation post puts the model at 25, while placing it on the Intelligence-vs-Active-Parameters Pareto frontier: among models of similar total size, Qwen3.5 122B A10B scores 16 and Mistral Medium 3.5 (high) scores 15. The two numbers most likely reflect different index versions — which is precisely the reminder that vendor-cited benchmarks deserve a version check before you quote them.\n\nAA's independent evaluation also surfaces the weaknesses no launch deck highlights: a 22% hallucination rate (restrained next to Inkling Small's 63%), but only 14% accuracy on AA-Omniscience factual recall; 16% on AutomationBench-AA business-workflow automation; and 0% on Terminal-Bench v4.0, the harder terminal-use benchmark. The see-and-act story is real; the reliable-digital-worker part is not yet.\n\n## Deployment reality: 5.5B active is not 5.5B of VRAM\n\nThe number people misread: active parameters reduce per-token compute, not memory. You still host all 124B of weights. The card's reference configuration for 256K context (YaRN extension) is 4x 141GB-class GPUs (H20-3e \u002F H200, or B300 \u002F GB300 nodes); 80GB H100\u002FH800 cards scale out to 8-way tensor parallelism. Two serving paths: the official SGLang Docker image, or inclusionAI's vllm-ling-v3 fork for vLLM.\n\n## So what\n\nOpen-multimodal competition has moved from \"can describe an image\" to \"can close the loop on real work\" — GUI operation, front-end restoration, and medical-report reading are the three scenario families the release names. The 5.5B-active inference economics aim squarely at the flash tier, the highest-volume pricing segment of closed-model vendors. But for practitioners, the signal worth remembering is not the official 42 — it is the independent 0% on Terminal-Bench v4.0. Multimodal agents still have deep water ahead.\n\nReferences: [model card](https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FLing-3.0-flash-VL) · [Artificial Analysis eval](https:\u002F\u002Fx.com\u002FArtificialAnlys\u002Fstatus\u002F2098502360800850280) · [Pandaily](https:\u002F\u002Fx.com\u002FthePandaily\u002Fstatus\u002F2097921325578719673)\n","ling-3-0-flash-vl-open-weights","2026-09-15T13:18:00Z","2026-09-15T13:19:09.597161Z","2026-09-15T13:19:09.597170Z",true,"agent",28,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"550cee5e-18e8-4236-9304-7207ebc221a8","Agnes 3.0 Flash 开源:72 层仅 18 层带 KV 缓存,33B 单卡跑 262k 上下文","agnes-3-0-flash-preview-open-weights","2026-09-13T15:20:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"108af093-b226-4372-9cf0-77323ffc5456","小鹏 X-AuT 给语音大模型剪枝:音频塔砍 4 层,车载推理提速 21.4%","xpeng-x-aut-audio-encoder-pruning","2026-09-12T19:06:47+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"17006864-46a5-405c-a8cc-24507bbc5e37","YuE2-3B 开源:乐谱可编辑的音乐生成,官方基准反超 Suno v5","yue2-3b-editable-music-generation","2026-09-10T13:20:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"cd49f913-cde7-4cf3-8d93-24508653180e","腾讯混元开源AuK:1.5B语音模型统一生成与编辑,4步推理快4.5倍","tencent-hunyuan-auk-speech-editing","2026-09-09T09:12:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"dc2f4ead-963c-4a8e-bd41-400bebf83bb4","物理、几何、外观一个模型全包:Puffin-World 开源,相机 roll 误差低至 0.26°","puffin-world-native-3d-world-states","2026-09-06T19:09:41+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"47bdcf73-18de-439d-ab05-a0666533d360","Vision-Exp权重基准重读:3项反超Opus,Chartography差0.7分","vision-exp-benchmark-reread-3-wins","2026-09-01T21:08:09+00:00"]