[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-hy4-tops-open-source-benchlm-september":3,"news-related-741bd34c-7134-4e8e-ab45-4f53dc576a6b":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"741bd34c-7134-4e8e-ab45-4f53dc576a6b","腾讯 Hy4 登顶 9 月开源榜:79.87 分超 Qwen3.8 Max,Anthropic 包揽总榜前三","BenchLM 9 月 1 日刷新榜单:腾讯 Hy4 preview 以 79.87 分成为最高分开源权重模型,反超 Qwen3.8 Max(79.4),位列 228 模型总榜第 6;Anthropic 三款 Claude 包揽总榜前三。","9 月 1 日,聚合基准平台 BenchLM 刷新了 9 月榜单。开源权重阵营的第一名换了人:腾讯 8 月 28 日发布的 Hy4 preview 以 79.87 分(满分 100)登上开源权重榜首,反超 Qwen3.8 Max 的 79.4 分,同时在 228 个已排名模型的总榜上位列第 6。上个月同一榜单的开源第一还是 Qwen3.8 Max(79.2 分),一个月内开源旗手完成交接。\n\n## 总榜格局:Anthropic 包揽前三\n\n先把总榜前几名摆出来:Claude Mythos 5(83.57)、Claude Fable 5(83.32)、Claude Opus 5(83.24)、GPT-5.6 Sol(82.39)、Kimi K3(80.78),然后才是 Hy4 preview。前五名在榜单上都不是开源权重模型,Hy4 是它们身后第一个开放权重的模型。按厂商平均分算,Anthropic 以 83.4 分(20 款模型)领先,OpenAI 76.5(38 款)、阿里巴巴 74.7(23 款)随后。BenchLM 的六个月发布记录显示 178 次发布、5 次领跑者更替,Hy4 preview 被标为 8 月发布的最高分模型。\n\n## Hy4 的分数是怎么来的\n\nBenchLM 的总分由 8 个加权类别构成,权重最高的是 Agentic(22%)和 Coding(20%)——恰好是 Hy4 最强的两块:Agentic 类别排名第 8(142 个模型中,95 百分位),Coding 排名第 12(147 个中,92 百分位)。\n\n细看单行基准,Hy4 有四个项目拿下全场已验证最佳:WideResearch 83.9%、JobBench 61.7%、BankerToolBench 78.6%、数学类 Apex 74.2%。知识类也不弱:GPQA 92.3%,Humanity's Last Exam(带工具)55.4%。编码侧 Terminal-Bench 2.1 拿到 85.4%,距离榜首 GLM-5.3 的 88.2% 只差 2.8 个百分点。\n\n工程侧,Hy4 preview 提供 1M 上下文窗口,腾讯在 Apache 2.0 许可下同时发布 BF16 检查点和 FP8 量化版,官方部署配方支持 vLLM 和 SGLang 自托管,但该检查点没有官方托管的 token 定价。\n\n## 三盆冷水\n\n第一,证据标签是 Estimated。Hy4 的档案页只有 28 行可展示的基准证据(平台共追踪 408 个基准槽位),Reasoning、Knowledge、Math、Multimodal 等类别都未达到排名门槛——平台明确说明这代表证据深度不足,不是能力为零,但选型时确实无从比较。\n\n第二,两个明显的短板项目。ProgramBench(从零重建程序)只有 17.5%,该项目最佳是 Claude Opus 5 的 93.0%,差距 75.5 个百分点;Agents' Last Exam 22.8%,Qwen3.8 Max 是 52.4%。SWE-bench Pro 65.7% 对 Claude Mythos 5 的 80.3% 也还有 14.6 个百分点的距离。\n\n第三,它叫 preview。速度未测、首 token 延迟未测、参数量未标注。从 4 月的 Hy3 Preview 到 7 月的 Hy3 再到现在,腾讯的 Hy 系列还在快速迭代期。\n\n## 怎么读这次换榜\n\n开源阵营的竞争焦点已经变了:不是「总分能不能逼近闭源」,而是在 Agentic 和 Coding 这两个权重最高、商业上最值钱的类别里正面竞争。Hy4 的登顶发生在 Agentic 权重 22% 的评价体系下,而且它的四个全场最佳全部来自 Agent 和工具使用类基准。\n\n对选型的人来说,这份榜单的正确用法不是看总分,而是点进单行证据:如果你的工作流是浏览器研究、金融工具调用,Hy4 的单行数据现在是全场最好;如果是长程程序重建,它目前还不行。排名是快照,证据才是决策依据——这话说给所有看到「登顶」两个字就准备冲进去的人。\n\n参考:BenchLM 9 月榜单(https:\u002F\u002Fbenchlm.ai\u002F)与 Hy4 preview 档案页(https:\u002F\u002Fbenchlm.ai\u002Fmodels\u002Fhy4-preview)。","https:\u002F\u002Fbenchlm.ai\u002F","34bc82d5-fd84-4d28-b5d4-4c057d78f972",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"7d58e50a-744d-4952-9761-fbb4bc1507f8","en","Tencent Hy4 Tops September Open-Weight Leaderboard at 79.87, Beating Qwen3.8 Max","BenchLM's September 1 refresh shows Tencent's Hy4 preview (released Aug 28) as the top open-weight model at 79.87\u002F100, overtaking Qwen3.8 Max at 79.4 and ranking #6 overall among 228 ranked models. Anthropic holds the top three overall spots.","On September 1, benchmark aggregation platform BenchLM refreshed its leaderboard. The open-weight crown changed hands: Tencent's Hy4 preview, released August 28, took the top open-weight spot with a score of 79.87 out of 100, overtaking Qwen3.8 Max at 79.4, and landing at #6 overall among 228 ranked models. A month earlier, Qwen3.8 Max (79.2) was still the open-weight leader on the same leaderboard — the flag has passed in a single month.\n\n## Overall Picture: Anthropic Sweeps the Top Three\n\nThe overall leaders: Claude Mythos 5 (83.57), Claude Fable 5 (83.32), Claude Opus 5 (83.24), GPT-5.6 Sol (82.39), and Kimi K3 (80.78), with Hy4 preview right behind. None of the top five are open-weight; Hy4 is the first open-weights model after them. By provider average, Anthropic leads at 83.4 (20 models), followed by OpenAI at 76.5 (38 models) and Alibaba at 74.7 (23 models). BenchLM's six-month release record shows 178 releases and 5 lead changes, with Hy4 preview marked as the highest-scoring model released in August.\n\n## Where Hy4's Score Comes From\n\nBenchLM's overall score is a weighted average across 8 categories, with Agentic (22%) and Coding (20%) carrying the most weight — exactly Hy4's two strongest areas: Agentic ranks #8 of 142 (95th percentile) and Coding ranks #12 of 147 (92nd percentile).\n\nAt the individual benchmark level, Hy4 holds the best verified results in four rows: WideResearch at 83.9%, JobBench at 61.7%, BankerToolBench at 78.6%, and Apex (math) at 74.2%. Knowledge holds up too: GPQA at 92.3% and Humanity's Last Exam with tools at 55.4%. On the coding side, Terminal-Bench 2.1 came in at 85.4%, just 2.8 points behind the leader GLM-5.3 at 88.2%.\n\nOn the engineering side, Hy4 preview offers a 1M context window. Tencent published both the BF16 checkpoint and a separate FP8 quantization under Apache 2.0, with official deployment recipes for self-hosting via vLLM or SGLang; no first-party hosted token pricing exists for this checkpoint.\n\n## Three Caveats\n\nFirst, the evidence label is Estimated. Hy4's profile shows only 28 source-displayable benchmark rows out of 408 tracked slots, and categories like Reasoning, Knowledge, Math, and Multimodal have not met the ranking threshold — the platform is explicit that this reflects evidence depth, not zero capability, but it does limit comparability.\n\nSecond, two visible weak spots. ProgramBench (rebuilding programs from scratch) sits at 17.5% versus a best verified 93.0% from Claude Opus 5 — a 75.5-point gap; Agents' Last Exam is 22.8% against Qwen3.8 Max's 52.4%. SWE-bench Pro at 65.7% also trails Claude Mythos 5's 80.3% by 14.6 points.\n\nThird, it's called a preview. Speed and time-to-first-token are unmeasured, and the parameter count is not yet sourced. From Hy3 Preview in April to Hy3 in July and now, Tencent's Hy line is still iterating fast.\n\n## How to Read This Change at the Top\n\nThe competition in open weights has shifted: it's no longer about whether the overall score can approach closed models, but about head-to-head competition in Agentic and Coding — the two highest-weighted and most commercially valuable categories. Hy4's rise happened under an evaluation system where Agentic carries 22% weight, and all four of its best-in-field results come from agent and tool-use benchmarks.\n\nFor anyone choosing models, the right way to use this leaderboard is not the overall score but the individual evidence rows: if your workflow is browser research or financial tool-calling, Hy4's row-level numbers are currently the best in the field; for long-horizon program rebuilding, it is not there yet. Rankings are snapshots; evidence is the decision basis — worth remembering before anyone charges in on the word tops.\n\nReference: BenchLM September leaderboard (https:\u002F\u002Fbenchlm.ai\u002F) and the Hy4 preview profile (https:\u002F\u002Fbenchlm.ai\u002Fmodels\u002Fhy4-preview).","tencent-hy4-tops-open-source-benchlm-september","2026-09-01T17:10:00Z","2026-09-01T17:10:09.036836Z","2026-09-01T17:10:09.036845Z",true,"agent",93,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"80315de0-7eb3-491a-b2e6-103a691a8bd7","Nanbeige4.2-3B 用 Looped Transformer 在 11 项基准上跑赢 Qwen3.5-9B","nanbeige-4-2-3b-looped-transformer-agentic-3b-beats-qwen3-5-9b","2026-07-30T10:30:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"6f1f105b-8e80-4b2c-b88c-b392556952aa","2026年本地LLM深度评测：开源模型性能全解析","local-llm-2026-deep-eval-swe-bench-aime","2026-04-25T11:15:00+00:00"]