[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen3-8-max-benchlm-aug-2026":3,"topics-all":38,"news-related-dcb1506b-87fa-422f-89e0-bb62afcc2b4c":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"dcb1506b-87fa-422f-89e0-bb62afcc2b4c","BenchLM 8 月榜:Qwen3.8 Max 79.2 分领跑开源 LLM,MiniMax M3 跻身三强","BenchLM 8 月 27 日开源 LLM 榜单:Qwen3.8 Max 以 79.2 分领跑(全榜第 6)、Qwen3.8-27B 72.6 分第二、MiniMax M3 68.6 分第三;12 个有真实部署记录的开源模型,中国厂商占 8 席,Apache\u002FMIT 占一半;开源模型首次稳定进入通用通道前 10。","开源大模型生态在 2026 年 8 月悄悄翻过一页:头部阵营第一次同时出现了 Qwen、阿里、Anthropic 之外的第三方厂商,头部模型与第二三名的差距也被拉开了 7 分以上。第三方榜单 BenchLM.ai 在 8 月 27 日校验的 Open-Source LLM Leaderboard 2026 给出了一份冷峻的快照:在前沿开源榜单上,阿里通义千问团队再次屠榜——Qwen3.8 Max 以 79.2 分位列第一,Qwen3.8-27B 以 72.6 分位列第二,MiniMax M3 以 68.6 分位列第三(原始数据:[benchlm.ai\u002Fbest\u002Fopen-source](https:\u002F\u002Fbenchlm.ai\u002Fbest\u002Fopen-source))。\n\n## 头部:千问系列霸榜,2.4T Max 第一次开源\n\nQwen3.8 Max 的 79.2 分比第二名 Qwen3.8-27B 高出 6.6 分,比第三名 MiniMax M3 高出 10.6 分。在 BenchAlign v5 公开打分通道上,这是开源与闭源阵营差距最小的一次——Qwen3.8 Max 在全榜(226 个模型)中排名第 6,在 105 个有源验证证据的子集中同样排名第 6。Qwen 团队在 OpenLM.ai 的官方公告里确认,Qwen3.8 Max 基于 Qwen3.5 架构扩展到 2.4 万亿参数,首次把 Max 级权重以 Apache\u002FMIT 方式开源(来源:[openlm.ai\u002Fqwen3.8](https:\u002F\u002Fopenlm.ai\u002Fqwen3.8))。\n\nQwen3.8-27B 走的是\"小而精\"路线——262K 上下文,72.6 分,适合单机单卡或双卡 H100 部署。它和 Qwen3.8 Max 一起,把开源模型在 BenchAlign v5 通用通道上的前两名都锁定在阿里手里。这也是 BenchLM 给出的第 12 名 MiniCPM5-1B(11.4 分)以来,103 个开源模型里最悬殊的头尾差距。\n\n## 第三:MiniMax M3 拿到 68.6 分,1M 上下文进入头部\n\nMiniMax M3 是 8 月开源榜上的一匹冷门黑马。68.6 分对应 1M 上下文窗口,显式推理模式支持,在 BenchLM 的子集证据标签里被标注为\"Supported\"——意味着分数来自可复核的公开源数据,不是估算。M3 的 90% 置信区间为 63.31–73.93,区间下沿跟 Dots Studio 的 dots3-note Preview(68.6)重合,这意味着开源阵营在 68 分附近的密度已经相当高。\n\n值得对比的是,M2.7 同厂版本仅 63.1 分。M3 相对 M2.7 提升约 5.5 分,代价是参数规模和推理硬件需求——官方数据指向 8× H100 起步(来源:[benchlm.ai\u002Fmodels\u002Fminimax-m3](https:\u002F\u002Fbenchlm.ai\u002Fmodels\u002Fminimax-m3))。\n\n## 部署视角:有\"实战记录\"的 12 个模型都是中国系或美国老厂\n\nBenchLM 的 8 月榜单里,只有 12 个开源模型进入了\"部署目录\"——意思是除了跑分,还有真实部署记录可以审计。许可证方面,MIT\u002FApache 2.0 阵营有 6 个:GLM-5.1(智谱,67 分)、Gemma 4 31B(Google,60.3)、DeepSeek-R1(51.1)、Qwen3.6-27B(53.7)、Mistral Small 4(46.4)、DeepSeek V3(44.5)。社区\u002F自定义许可证的有 6 个:Kimi K2.6\u002FK2.5\u002FK2.7 Code、Qwen2.5-72B、Llama 4 Scout\u002FMaverick。\n\n这条部署目录透露两个信号。第一,中国厂商占了 8 席(智谱、阿里、Kimi、DeepSeek),美国厂商 4 席(Google、NVIDIA、Mistral、Meta)。第二,真正能在单机单卡跑的开源旗舰只剩下 Gemma 4 31B、Qwen3.6-27B 这种 27B–31B 区间的——再往上 Qwen3.8 Max、GLM-5.1、Kimi K2.6 都需要 8× H100 的 640GB 总显存(原始数据同上,见 BenchLM 部署目录)。\n\n## 几个被反复引用的\"中国实验室超大参数\"在这份榜里怎么样\n\nSolidot 在 8 月 17 日转述的 Hugging Face 报告提到:中国发布的 178 个 200 亿以上参数模型里,55% 用 Apache 2.0、22% 用 MIT,但大部分有非商业使用限制。NVIDIA Nemotron 3 Ultra(5610 亿参数)和 Thinking Machines 的 Inkling(9520 亿参数,基于中国模型构建)也在该报告中被点名。在 BenchLM 的开源榜上,Nemotron 3 Ultra 仅排第 66 位(46.4 分),Inkling 排第 9(66.9 分),Inkling-Small 排第 11(63.7 分)。参数规模并没有直接转化为榜单位置——Qwen3.8-27B(27B 密集)用 1\u002F90 的参数量把 9520 亿参数的 Inkling 压在身后。\n\n## 行业影响:开源模型在通用通道上第一次摸到全榜前 10\n\nBenchAlign v5 通用通道是开源\u002F闭源同台打分的通道。在 8 月榜单里,Qwen3.8 Max 已经是全榜(226 个模型含闭源)的第 6 名——这是开源模型在这条通道上第一次稳定出现在前 10。对比 7 月 14 日的快照,开源榜首 MiniMax M3 的 79.2 分(同分数是 Qwen3.8 Max 在 8 月拿到的)当时对应的是全榜第 4 左右。8 月的开源榜单变化意味着:头部闭源模型在通用通道上的领先幅度正在收敛,而收敛的主力不是 Meta、不是 Mistral,是阿里的 Qwen 系列。\n\n所以这条 8 月榜单讲的不是\"哪家最强\",而是\"开源模型第一次在通用基准上不输闭源头部,而且领跑的不是西方实验室\"。下个月榜单如果 Qwen3.9 系列或 GLM-6 出现在前 3,这条趋势就算坐实了。\n\n来源汇总:\n- BenchLM 8 月 27 日榜单:[benchlm.ai\u002Fbest\u002Fopen-source](https:\u002F\u002Fbenchlm.ai\u002Fbest\u002Fopen-source)\n- Qwen3.8 Max 模型卡:[benchlm.ai\u002Fmodels\u002Fqwen3-8-max](https:\u002F\u002Fbenchlm.ai\u002Fmodels\u002Fqwen3-8-max)\n- Qwen 官方公告:[openlm.ai\u002Fqwen3.8](https:\u002F\u002Fopenlm.ai\u002Fqwen3.8)\n- Solidot 8 月 17 日转述 Hugging Face 报告:[solidot.org\u002Fstory?sid=85118](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85118)","https:\u002F\u002Fbenchlm.ai\u002Fbest\u002Fopen-source","34bc82d5-fd84-4d28-b5d4-4c057d78f972",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1b66d2e4-deec-4406-acc5-8b0a939638e7","en","BenchLM August 2026: Qwen3.8 Max Leads Open-Source LLMs at 79.2, MiniMax M3 Cracks the Top Three","BenchLM's August 27 leaderboard for open-weight LLMs: Qwen3.8 Max leads at 79.2 (rank 6 overall), Qwen3.8-27B is second at 72.6, MiniMax M3 is third at 68.6. Of the 12 models with real deployment records, Chinese vendors hold eight seats and Apache\u002FMIT licenses cover half — the first time open-weight models have stably entered the top 10 of the unified public benchmark.","The open-weight LLM landscape quietly turned a page in August 2026: the leaderboard now shows a non-Qwen, non-Anthropic third-party model in the top tier for the first time, and the gap between first place and the rest has stretched past seven points. The third-party leaderboard BenchLM.ai, verified on August 27, 2026, published its Open-Source LLM Leaderboard 2026 with a stark snapshot. On the frontier open-weight board, Alibaba's Qwen team swept the podium: Qwen3.8 Max at 79.2, Qwen3.8-27B at 72.6, and MiniMax M3 at 68.6 (source: benchlm.ai\u002Fbest\u002Fopen-source).\n\n## The Top: Qwen Sweeps, 2.4T Max Goes Open-Source for the First Time\n\nQwen3.8 Max's 79.2 beats Qwen3.8-27B by 6.6 points and MiniMax M3 by 10.6 points. On the public BenchAlign v5 lane this is the narrowest gap yet between open and closed tiers. Qwen3.8 Max ranks #6 out of 226 models on the full board, and #6 out of 105 source-verified models. Qwen's official announcement on OpenLM.ai confirms that Qwen3.8 Max extends the Qwen3.5 architecture to 2.4 trillion parameters and releases Max-class weights under Apache\u002FMIT for the first time (source: openlm.ai\u002Fqwen3.8).\n\nQwen3.8-27B takes the \"small but precise\" lane — 262K context, 72.6 points, deployable on a single H100 or a dual-GPU rig. Together with Qwen3.8 Max it locks the top two of the BenchAlign v5 general lane for Alibaba. This is also the widest head-to-tail spread on the 103-model list, ending at MiniCPM5-1B at 11.4.\n\n## Third Place: MiniMax M3 Hits 68.6 With 1M Context\n\nMiniMax M3 is the dark horse of the August open-weight board. Its 68.6 pairs with a 1M context window and an explicit reasoning mode, and BenchLM tags its evidence as \"Supported\", meaning the score comes from auditable public data rather than estimation. M3's 90% confidence interval is 63.31–73.93; the lower bound overlaps with Dots Studio's dots3-note Preview (68.6), which means open-weight density around 68 is now tight.\n\nFor comparison, M2.7 from the same vendor scored 63.1. M3 improves by about 5.5 points over M2.7, at the cost of parameter count and inference hardware — official specs point to an 8× H100 floor (source: benchlm.ai\u002Fmodels\u002Fminimax-m3).\n\n## Deployment Lens: 12 Models With Real-World Records Are All Chinese Vendors or US Incumbents\n\nBenchLM's August board lists only 12 open-weight models in its deployment catalog — meaning they have auditable real-world deployment records beyond leaderboard scores. On the license side, MIT\u002FApache 2.0 covers six seats: GLM-5.1 (Zhipu, 67), Gemma 4 31B (Google, 60.3), DeepSeek-R1 (51.1), Qwen3.6-27B (53.7), Mistral Small 4 (46.4), and DeepSeek V3 (44.5). Community or custom licenses cover the remaining six: Kimi K2.6\u002FK2.5\u002FK2.7 Code, Qwen2.5-72B, and Llama 4 Scout\u002FMaverick.\n\nThe deployment catalog tells two stories. First, Chinese vendors hold eight seats (Zhipu, Alibaba, Kimi, DeepSeek), US vendors four (Google, NVIDIA, Mistral, Meta). Second, the only frontier open-weight models you can actually run on a single consumer GPU are the 27B–31B class — Gemma 4 31B and Qwen3.6-27B. Everything heavier, including Qwen3.8 Max, GLM-5.1, and Kimi K2.6, demands 8× H100 at 640GB total VRAM.\n\n## What Happens to the \"Chinese Lab Giant-Parameter\" Talking Point\n\nSolidot's August 17 relay of the Hugging Face report cited that 178 Chinese models above 20B parameters used Apache 2.0 (55%) or MIT (22%) — though most still carry non-commercial restrictions. The same report named NVIDIA's Nemotron 3 Ultra (561B) and Thinking Machines' Inkling (9520B, built on a Chinese model). On BenchLM's open-weight board, Nemotron 3 Ultra sits at rank 66 (46.4 points), Inkling at rank 9 (66.9), Inkling-Small at rank 11 (63.7). Parameter count does not translate directly into board position: Qwen3.8-27B (27B dense) uses 1\u002F90 the parameters to beat 9520B Inkling.\n\n## Industry Impact: Open-Weight Models Touch the Top 10 on the General Lane for the First Time\n\nBenchAlign v5 is a unified lane where open and closed models are scored side by side. On the August board, Qwen3.8 Max is rank 6 on the full 226-model board including closed models — the first time an open-weight model has stably landed in the top 10 on this lane. Compared to the July 14 snapshot, the same 79.2 score held by MiniMax M3 then corresponded to roughly rank 4 on the full board. The August shift means the closed-tier lead is compressing, and the compression is being driven not by Meta or Mistral but by Alibaba's Qwen series.\n\nSo the August board is not about \"who is strongest\", it is about \"open-weight models can finally match the closed-tier head on a general benchmark, and the leader is not a Western lab\". If Qwen3.9 or GLM-6 lands in the top three next month, this trend is no longer a fluke.\n\nSources:\n- BenchLM August 27 leaderboard: benchlm.ai\u002Fbest\u002Fopen-source\n- Qwen3.8 Max model card: benchlm.ai\u002Fmodels\u002Fqwen3-8-max\n- Qwen official announcement: openlm.ai\u002Fqwen3.8\n- Solidot August 17 relay of Hugging Face report: solidot.org\u002Fstory?sid=85118","qwen3-8-max-benchlm-aug-2026","2026-08-28T06:00:00Z","2026-08-28T01:06:19.257925Z","2026-08-28T01:06:19.257937Z",true,"agent",420,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"4c9f74d4-0252-4e86-8b6e-85d38788eea6","开源编程模型三选一:GLM-5.2、DeepSeek V4、Qwen3.6","glm-5-2-deepseek-v4-qwen-3-6-coding","2026-07-27T06:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"2e27016d-b90e-45c7-825a-41fd1e435c80","JHU 新研究:组合持续学习机制,百任务记忆留存从 1.2% 提到 34.9%","compose-cl-long-horizon-memorization","2026-09-16T15:10:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"2731ed1c-17c3-4d85-9174-983cf50743e3","地铁售票机上的 AI 大考:2.6GB 端侧模型 91.32 分超 GPT-5.6,规则基线也拿 84.6","metrollm-bench-transit-kiosk-llm","2026-09-12T23:08:18+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"54b86d93-0fd0-4107-9353-9b79a1446f69","NVIDIA 开源 IMO 金牌完整配方:30\u002F42 分、561B 双专家、算力账本全公开","nvidia-nemotron-imo-gold-open-recipe","2026-09-11T17:13:27+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"d41175a7-ad10-4e00-9017-a148fa0a77b3","BenchMIRT 把 LLM 基准拆到单题:Ai2 想让模型排名不再「一张考卷定生死」","ai2-benchmirt-llm-benchmark-audit","2026-09-10T11:05:05+00:00"]