On September 1, benchmark aggregation platform BenchLM refreshed its leaderboard. The open-weight crown changed hands: Tencent's Hy4 preview, released August 28, took the top open-weight spot with a score of 79.87 out of 100, overtaking Qwen3.8 Max at 79.4, and landing at #6 overall among 228 ranked models. A month earlier, Qwen3.8 Max (79.2) was still the open-weight leader on the same leaderboard — the flag has passed in a single month.
Overall Picture: Anthropic Sweeps the Top Three
The overall leaders: Claude Mythos 5 (83.57), Claude Fable 5 (83.32), Claude Opus 5 (83.24), GPT-5.6 Sol (82.39), and Kimi K3 (80.78), with Hy4 preview right behind. None of the top five are open-weight; Hy4 is the first open-weights model after them. By provider average, Anthropic leads at 83.4 (20 models), followed by OpenAI at 76.5 (38 models) and Alibaba at 74.7 (23 models). BenchLM's six-month release record shows 178 releases and 5 lead changes, with Hy4 preview marked as the highest-scoring model released in August.
Where Hy4's Score Comes From
BenchLM's overall score is a weighted average across 8 categories, with Agentic (22%) and Coding (20%) carrying the most weight — exactly Hy4's two strongest areas: Agentic ranks #8 of 142 (95th percentile) and Coding ranks #12 of 147 (92nd percentile).
At the individual benchmark level, Hy4 holds the best verified results in four rows: WideResearch at 83.9%, JobBench at 61.7%, BankerToolBench at 78.6%, and Apex (math) at 74.2%. Knowledge holds up too: GPQA at 92.3% and Humanity's Last Exam with tools at 55.4%. On the coding side, Terminal-Bench 2.1 came in at 85.4%, just 2.8 points behind the leader GLM-5.3 at 88.2%.
On the engineering side, Hy4 preview offers a 1M context window. Tencent published both the BF16 checkpoint and a separate FP8 quantization under Apache 2.0, with official deployment recipes for self-hosting via vLLM or SGLang; no first-party hosted token pricing exists for this checkpoint.
Three Caveats
First, the evidence label is Estimated. Hy4's profile shows only 28 source-displayable benchmark rows out of 408 tracked slots, and categories like Reasoning, Knowledge, Math, and Multimodal have not met the ranking threshold — the platform is explicit that this reflects evidence depth, not zero capability, but it does limit comparability.
Second, two visible weak spots. ProgramBench (rebuilding programs from scratch) sits at 17.5% versus a best verified 93.0% from Claude Opus 5 — a 75.5-point gap; Agents' Last Exam is 22.8% against Qwen3.8 Max's 52.4%. SWE-bench Pro at 65.7% also trails Claude Mythos 5's 80.3% by 14.6 points.
Third, it's called a preview. Speed and time-to-first-token are unmeasured, and the parameter count is not yet sourced. From Hy3 Preview in April to Hy3 in July and now, Tencent's Hy line is still iterating fast.
How to Read This Change at the Top
The competition in open weights has shifted: it's no longer about whether the overall score can approach closed models, but about head-to-head competition in Agentic and Coding — the two highest-weighted and most commercially valuable categories. Hy4's rise happened under an evaluation system where Agentic carries 22% weight, and all four of its best-in-field results come from agent and tool-use benchmarks.
For anyone choosing models, the right way to use this leaderboard is not the overall score but the individual evidence rows: if your workflow is browser research or financial tool-calling, Hy4's row-level numbers are currently the best in the field; for long-horizon program rebuilding, it is not there yet. Rankings are snapshots; evidence is the decision basis — worth remembering before anyone charges in on the word tops.
Reference: BenchLM September leaderboard (https://benchlm.ai/) and the Hy4 preview profile (https://benchlm.ai/models/hy4-preview).