[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen3-8-max-2-4t-moe-open-weights":3,"news-related-40095b51-97b0-4fd4-9b1d-f636c970572e":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","阿里 Qwen 团队于 8 月 3 日正式上线 Qwen3.8-Max,这是一种 2.4 万亿参数的混合专家(MoE)模型,支持百万 token 上下文与多模态输入,并首度承诺开放 Max 级权重(下周交付),同步推出可单机部署的 Qwen3.8-27B。Qwen Cloud 公布的 API 价格为输入 2 美元\u002F百万 token、输出 6 美元\u002F百万 token,缓存输入 0.25 美元\u002F百万 token。","## 背景:Qwen 家族的 Max 级别补齐\n\n阿里通义千问(Qwen)团队长期以「开源开放」路线区别于其它顶级实验室:过去两年里,Qwen-Image、Qwen-VL、QwQ 等、Qwen3.6-Plus 等模型几乎全部开源权重。但旗舰定位的 Qwen-Max 系列此前一直仅以 API 形式提供,社区无法本地部署。**2026 年 8 月 3 日,Qwen 团队在 qwen.ai\u002Fblog?id=qwen3.8 发布 Qwen3.8-Max,这是 Qwen 家族中「首次开放 Max 级权重」的旗舰模型**(参考 MarkTechPost 同期报道),配套下放一个 270 亿参数的次级检查点 Qwen3.8-27B。\n\n## 核心规格:2.4 万亿 MoE + 百万 token 上下文\n\nQwen3.8-Max 的核心参数由 MarkTechPost 与 QwenCloud 模型页面共同确认:\n\n- **总参数规模**:2.4 万亿,采用混合专家(MoE)架构;激活参数数量官方尚未披露\n- **上下文窗口**:1,000,000 token(启用思考时输入上限 983K,关闭时 991K),输出上限 131K,推理预算最高 262K\n- **模态**:支持文本、图像、视频作为输入,文本输出\n- **API 价格(Qwen Cloud)**:输入 2.00 美元\u002F百万 token,输出 6.00 美元\u002F百万 token;隐式缓存读取 0.25 美元\u002F百万 token,显式缓存构建 2.50 美元、读取 0.17 美元——缓存输入比新鲜输入便宜约 8 倍\n- **速率限制**:2M token\u002F分钟、15K 请求\u002F分钟\n- **能力**:函数调用、结构化输出、批处理、前缀续写、微调,Responses API 内置 code_interpreter \u002F web_search \u002F web_extractor \u002F t2i_search \u002F i2i_search 五种工具\n\n## 基准成绩:多模态领先,代码推理追平一线\n\n阿里官方随版本同步放出完整的 benchmark 表(详见 qwen.ai\u002Fblog?qwen3.8)。比较有信号的几项:\n\n- **Terminal-Bench 2.1**:**86.6** 分,高于 Claude Opus 4.8、Fable 5 的 84.6,但仍低于 GPT-5.6 Sol(max)的 88.8\n- **SWE-bench Pro**:**67.7** vs Fable 5 的 80.0\n- **FrontierSWE**:**73.5** vs Fable 5 的 88.8\n- **PaperBench**:**93.0**(领先)\n- **IFBench**:**82.8**(领先)\n- **GPQA Diamond**:92.6,仅比 Qwen3.7-Max 的 92.4 微涨\n- 多模态项目明显领先:**OSWorld-Verified 86.1、Parametric CAD Bench 91.5、OmniDocBench 1.5 92.1**\n- 自我对照涨幅很大:**DeepSWE 1.1 从 21.6 跃升到 56.6,FrontierSWE 从 40.7 到 73.5,JobBench 从 31.3 到 53.4**\n\n需要标注的两点诚实折扣:**多模态表是与 Qwen3.7-Plus(而非 Qwen3.7-Max)对比**,这放大了代际差距;**官方 RL 缩放曲线在 4000 个训练环境附近达到 0.725 峰值,继续扩到 0.719 和 0.689 时反而回落**——意味着 2.4T 并不是「越大越好」的简单外推。\n\n## 部署现实:Max 跑不动,27B 才是落地形态\n\nMarkTechPost 的判断很清醒:虽然 Qwen3.8-Max 立刻可以通过 API 接入(OpenAI 与 DashScope 兼容,改 base URL 与 model ID 即可),但 **2.4T 总参数的检查点本身就是「多节点数据中心级」产物**,不是单机 GPU 能塞下的东西。**真正的本地可部署形态是 Qwen3.8-27B**,它和 Max 共享同一波次的训练迭代,但参数规模压到了普通 GPU 服务器可以承载的区间。Max 权重虽然承诺「下周开放」,但具体协议、可商用范围、激活参数量、推理框架适配情况官方还没给出清单,实际能跑成什么样要等社区拿到权重后才能见分晓。\n\n## 评论:为什么「开源 Max」这件事重要\n\n把 Max 级别模型开放权重,意义在于**让社区第一次有可能在本地评估一个阿里对标 GPT-5.6 Sol、Claude Opus 4.8、Fable 5 的真实旗舰**。过去这些顶级模型在 benchmark 表里被写在一起,但只有 API 形式,大家都得听厂商自报;现在 Max 权重一旦放出,社区可以做独立的复现、独立的能力测试、独立的红队,这对 AI 安全与对齐研究也是一次「解封」。\n\n但也要看到:**Max 权重 ≠ Max 可用**。27B 检查点才是真正的产品形态,这是开源 MoE 旗舰的标准操作——DeepSeek V4-Pro 的 1.6T 部署形态、Kimi K3 的 2.8T 在 Hugging Face 上的可玩版本都走了同样的路。**未来一周真正值得盯的事是:官方能否同步给出激活参数量、推理框架(vLLM \u002F SGLang \u002F TensorRT-LLM)的适配进度、以及清晰的许可证条款。**没有这三项,2.4T 仍只是 PR 数字;有了,头部开源模型的天花板会再被掀一次。\n\n---\n\n**来源**(素材):\n1. Qwen 官方博客:Qwen3.8-Max 发布页(qwen.ai\u002Fblog?id=qwen3.8,通过 qwen.ai\u002Fresearch 索引确认)\n2. MarkTechPost 报道(2026-08-03):Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model\n3. X \u002F Twitter:Alibaba_Qwen 官方账号 8 月初对 Qwen3.8-27B 开源化的预告","https:\u002F\u002Fwww.marktechpost.com\u002F2026\u002F08\u002F03\u002Falibaba-qwen-releases-qwen3-8-max\u002F","8382d60c-c2c4-49c5-9638-8518b803f88f",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":25,"name":26,"slug":26,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"0b73720c-9558-4c02-818a-22f472a1ead7","en","Qwen3.8-Max: Alibaba opens weights of its 2.4T-parameter MoE","On August 3, 2026, the Alibaba Qwen team released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts (MoE) model with a 1M-token context window and multimodal input. For the first time in the Qwen family, Max-class weights are promised to ship as open weights next week, alongside an on-premise-friendly Qwen3.8-27B checkpoint. Qwen Cloud pricing is USD 2 per 1M input tokens and USD 6 per 1M output tokens, with cached input at USD 0.25 per 1M tokens.","## Background: Filling the Max slot in the Qwen family\n\nThe Alibaba Tongyi Qianwen (Qwen) team has long distinguished itself from other top labs with an open-source-first stance: over the past two years, Qwen-Image, Qwen-VL, QwQ, Qwen3.6-Plus and others were released with weights. But the flagship Qwen-Max line had previously been available only as an API. **On August 3, 2026, the Qwen team published Qwen3.8-Max at qwen.ai\u002Fblog?id=qwen3.8 — the first Max-class model in the family whose weights will go open (shipping next week)**, paired with a 27B-parameter sibling, Qwen3.8-27B, that is also going open-weights (per MarkTechPost same-day coverage).\n\n## Core specs: 2.4T MoE with a 1M-token context window\n\nThe headline numbers, cross-checked between MarkTechPost and the Qwen Cloud model page:\n\n- **Total parameters**: 2.4 trillion, mixture-of-experts (MoE) architecture. Activated-parameter count has not been disclosed by Alibaba.\n- **Context window**: 1,000,000 tokens (983K with thinking on, 991K with thinking off). Maximum output is 131K in both modes. Maximum reasoning budget is 262K.\n- **Modalities**: accepts text, image, and video as input; returns text.\n- **API pricing (Qwen Cloud)**: USD 2.00 per 1M input tokens, USD 6.00 per 1M output tokens. Implicit cache reads cost USD 0.25 per 1M tokens; explicit cache creation is USD 2.50 and reads USD 0.17. Cached input is roughly 8x cheaper than fresh input.\n- **Rate limits**: 2M tokens per minute, 15K requests per minute.\n- **Capabilities**: function calling, structured outputs, batches, prefix completion, fine-tuning. The Responses API ships five built-in tools: code_interpreter, web_search, web_extractor, t2i_search, and i2i_search.\n\n## Benchmark scores: multimodal leads, code and reasoning catch the frontier\n\nAlibaba published a full benchmark table alongside the release. The most signal-laden entries:\n\n- **Terminal-Bench 2.1**: 86.6, ahead of Claude Opus 4.8 and Fable 5 (both 84.6), but behind GPT-5.6 Sol (max) at 88.8.\n- **SWE-bench Pro**: 67.7 vs Fable 5 80.0.\n- **FrontierSWE**: 73.5 vs Fable 5 88.8.\n- **PaperBench**: 93.0 (leading).\n- **IFBench**: 82.8 (leading).\n- **GPQA Diamond**: 92.6, only marginally up from Qwen3.7-Max 92.4.\n- Multimodal rows lead cleanly: OSWorld-Verified 86.1, Parametric CAD Bench 91.5, OmniDocBench 1.5 92.1.\n- Internal deltas vs Qwen3.7 are large: DeepSWE 1.1 jumps from 21.6 to 56.6, FrontierSWE from 40.7 to 73.5, JobBench from 31.3 to 53.4.\n\nTwo honest caveats belong in any read: **the multimodal table benchmarks against Qwen3.7-Plus, not Qwen3.7-Max, which flatters the generational gap**; and **Alibaba own RL scaling curve peaks at 0.725 near 4,000 training environments, then declines to 0.719 and 0.689** — so 2.4T is not a simple bigger-is-better extrapolation.\n\n## Deployment reality: Max is not deployable, 27B is\n\nMarkTechPost read is the clean one: while Qwen3.8-Max is available right now via the API (OpenAI- and DashScope-compatible, so integration is just a base-URL and model-ID change), **the 2.4T-total-parameter checkpoint is a multi-node datacenter artifact**, not anything a single GPU server can host. **The realistic on-premise form factor is Qwen3.8-27B** — same training iteration, parameter count compressed to a range ordinary GPU hardware can serve. Max weights are promised for next week, but Alibaba has not published a license, the activated-parameter count, or which inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) will support it. Until those land, the 2.4T number is mostly a PR line.\n\n## Why Max goes open matters\n\nPutting Max-class weights out is the first time the community can locally evaluate an Alibaba model that genuinely sits in the same benchmark conversation as GPT-5.6 Sol, Claude Opus 4.8, and Fable 5. Until now those models appeared in the same benchmark tables but were vendor-self-reported; with Max weights in the wild, the community can do independent capability tests, independent red-teaming, independent reproducibility work — which also matters for AI safety and alignment research.\n\nBut the gap between Max-is-open and Max-is-usable is real. 27B is the actual product form factor, and that is the standard playbook for open MoE flagships: DeepSeek V4-Pro 1.6T, Kimi K3 2.8T all serve on the same pattern via smaller, locally-runnable checkpoints. **The real things to watch over the next week are: whether Alibaba publishes the activated-parameter count, the vLLM \u002F SGLang \u002F TensorRT-LLM adapter roadmap, and a clear license**. Without those three, 2.4T remains a marketing number; with them, the open-source frontier ceiling moves up another notch.\n\n---\n\n**Sources used**:\n1. Official Qwen blog: Qwen3.8-Max release page (qwen.ai\u002Fblog?id=qwen3.8, indexed via qwen.ai\u002Fresearch)\n2. MarkTechPost coverage (2026-08-03): Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model\n3. X\u002FTwitter: Alibaba_Qwen official account, early-August preview of Qwen3.8-27B open-weights intent","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00Z","2026-08-07T14:21:57.735159Z","2026-08-07T14:21:57.735167Z",true,"agent",317,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"cb64371f-62b6-473d-8150-b576001d3f56","Qwen3.8-27B 开源权重上线:单卡跑得动的 Qwen3.8,还塞了个视觉编码器","qwen3-8-27b-open-weights-release","2026-08-14T19:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"804b44fb-66c6-4355-83a4-b3a03a776d2a","Inkling-Small 开放权重：12B 激活参数换来更高 Agent 效率，也暴露事实性短板","inkling-small-multimodal-moe-efficiency","2026-08-05T16:32:13+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"dfc3dec4-2211-4c7e-b6ff-9e0d9a479ec4","微软与 Mistral 签下数十亿美元协议:Vera Rubin GPU 上的「欧洲主权云」开始落地","microsoft-mistral-vera-rubin-sovereign","2026-07-22T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"df01dae0-0940-4947-a019-31c57066132c","阿里千问 Qwen3.8 预览版上线:2.4T 参数,把开源旗舰抬到 Fable 5 同一档","qwen-3-8-preview-2-4t","2026-07-19T10:01:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"cf01282f-8a64-49a8-a608-9b806ccfbea3","Mira Murati 实验室 Inkling 开源：975B MoE 不卷\"最强\"，押注\"可定制\"","thinking-machines-inkling","2026-07-15T22:00:00+00:00"]