[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ai2-benchmirt-llm-benchmark-audit":3,"topics-all":38,"news-related-d41175a7-ad10-4e00-9017-a148fa0a77b3":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"d41175a7-ad10-4e00-9017-a148fa0a77b3","BenchMIRT 把 LLM 基准拆到单题:Ai2 想让模型排名不再「一张考卷定生死」","Ai2 推出多维 IRT 方法 BenchMIRT，把 16 个 LLM 评测基准的 3.4 万道题拆开分析，发现 BBQ、WMDP 等榜单的真实信号和「标签说的能力」并不一致，模型排名可能被一两个隐性维度主导。","Ai2（Allen Institute for AI）在 Hugging Face 博客公开了 BenchMIRT——一种多维 IRT 方法。它把 LLM 基准拆到单题：16 个基准、3.4 万道题、100 个开源模型，分别给出 safety 和 general reasoning 两个维度的能力估计。结果很尴尬——大量被默认归到「安全类」的基准，跑出来的真实信号其实更接近「通用推理」。\n\n## 为什么这件事值得专门写\n\n过去两年，LLM 排行榜几乎被 MMLU、GPQA、Chatbot Arena 一统天下。同一个「第一名」在不同榜单上动辄差 10 分，OpenAI、Anthropic、Google 三家互有胜负。根问题是现行榜单把能力混着测——一个 BBQ 题目（测社会偏见）同时需要「理解谁是爷爷、谁是孙子」+「会不会写出刻板印象」，两个因子搅在一起，低分到底是 safety 还是阅读理解？BenchMIRT 想回答的就是这种问题。\n\n## BenchMIRT 怎么做、发现了什么\n\n方法不新：单维 IRT 在 SAT、GRE 用了几十年，原理是根据答题情况反推每道题的难度和区分度。Ai2 之前 Fluid Benchmarking 引进过单维版，这次升级到多维。实验规模：100 个开源 LLM × 16 个基准 × 3.4 万题。6 个测通用推理（MMLU-Pro、GPQA、MATH、BBH）；10 个来自 Ai2 自家 Olmo 3 安全套件（HarmBench、StrongReject、WildJailbreak、BBQ、WMDP、XSTest）。\n\n关键发现：BenchMIRT 没被告知「哪些基准测的是 safety 还是 reasoning」，却自己稳定恢复出两个主维度。\n\n几条反直觉结论：BBQ（测社会偏见）和「通用推理」的相关性远强于「安全」；WMDP（测危险知识）方向相反——推理越强的模型越会「安全地装作不知道」，被 WMDP 判「答错」；HarmBench 的版权题跑出来更像通用推理维度，和有害请求类不在同一信号上。\n\n## 对 LLM 评测行业的实际影响\n\n「基准瘦身」变得可执行。只保留 10% 的题目，模型排序基本不变；保留 50% 时排序几乎完全一致。MMLU-Pro 等被诟「刷不动」的榜可压缩到十分之一成本，效果损失很小。\n\n榜单设计者要重新审视「这个基准到底测什么」。BenchMIRT 预测单题答对与否的准确率是 79%，简单基线只有 70%。看似只提升 9 个点，但能让评测人员精确知道「哪几道题最能反映 safety 维度」。\n\n副作用也明显：BenchMIRT 给出的「哪些题最有区分度」，反过来就是「删掉哪些题就能让安全差的模型蒙混过关」。Ai2 坦承这是真实风险，但认为可解释性收益大于被绕过风险。\n\n## 几点观察\n\n100 个模型的训练截止时间都被锁在 2025 年 3 月之前，意味着 BenchMIRT 对 GPT-5 系列、Claude Opus 5、Gemini 3.x 这些 2025 下半年到 2026 年的新模型不一定成立。\n\n另外，BenchMIRT 找到的维度强烈依赖喂进去的基准池。这次偏 safety + reasoning，下次喂代码、数学、多模态，出来的大概率会变成「代码 \u002F 数学 \u002F 多模态」三足鼎立。**这套工具能告诉你「你现在测的能力到底有几种」，但不能告诉你「还有哪些能力是你没测的」**——这是它的盲区，也是未来 LLM 评测要补的一课。\n\n回到行业视角：Chatbot Arena 已是事实标准「全民榜」，但 ELO 只有总分。BenchMIRT 这类方法若能整合进主流榜单——给每个模型出一张「能力向量图」——LLM 评测可能从「谁的分高」升级到「谁在哪个能力上更强」。这件事比模型本身一次升级更有结构性意义。\n\n参考资料：[Ai2 在 Hugging Face 的官方博客](https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fallenai\u002Fbenchmirt)、[BenchMIRT 数据集](https:\u002F\u002Fhuggingface.co\u002Fcollections\u002Fallenai\u002Fbenchmirt)、[GitHub 代码](https:\u002F\u002Fgithub.com\u002Fallenai\u002FBenchMIRT)。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fallenai\u002Fbenchmirt","446fa1f9-89de-415a-a793-275018ca9899",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"ec12ed26-fe32-425b-ba93-ed8412307666","en","BenchMIRT Audits LLM Benchmarks Prompt by Prompt: Ai2 Wants Model Rankings Off a Single Test","Ai2 has open-sourced BenchMIRT, a multidimensional IRT method that decomposes 16 LLM benchmarks and 34,000 prompts across 100 open-weight models, finding that real signals in BBQ, WMDP and others often diverge from their stated capabilities.","Ai2 (Allen Institute for AI) has published BenchMIRT on the Hugging Face blog, a multidimensional IRT method that takes LLM benchmarks apart at the prompt level: 16 benchmarks, 34,000 prompts, 100 open-weight models, producing separate ability estimates along safety and general reasoning dimensions. The result is awkward: many benchmarks that get bucketed under \"safety\" actually carry a signal much closer to \"general reasoning\".\n\n## Why this is worth its own writeup\n\nFor the past two years, the LLM leaderboard scene has been dominated by MMLU, GPQA and Chatbot Arena. The same \"rank-one\" model can swing by 10 points between boards, and OpenAI, Anthropic and Google keep trading the top spot. The root cause is that today's benchmarks mix capabilities together. A BBQ question (testing social bias) needs both \"track who is the grandfather, who is the grandson\" and \"avoid stereotypical wording\"; the two factors get tangled, so a low score is hard to read as safety versus reading comprehension. BenchMIRT is built to answer exactly that question.\n\n## How BenchMIRT works, and what it found\n\nThe method is not new: single-dimensional IRT has been used in SAT and GRE for decades. The principle is to infer, from how a model answers, each prompt's difficulty and discriminative power. Ai2's earlier work, Fluid Benchmarking, brought single-dimensional IRT into LLM evaluation; this time the team moves to multi-dimensional IRT, where each benchmark can map onto multiple capabilities. The experiment covers 100 open-weight LLMs × 16 benchmarks × more than 34,000 prompts. Six benchmarks measure general reasoning (MMLU-Pro, GPQA, MATH, BBH), and ten come from Ai2's own Olmo 3 safety suite (HarmBench, StrongReject, WildJailbreak, BBQ, WMDP, XSTest).\n\nThe key finding: BenchMIRT is never told which benchmarks measure safety versus reasoning, yet it stably recovers two dominant dimensions on its own.\n\nA few counterintuitive results: BBQ's correlation with general reasoning is far stronger than with safety; WMDP (dangerous-knowledge benchmark) points the other way — stronger reasoning models tend to \"safely pretend they don't know\" and get marked wrong by WMDP; HarmBench's copyright prompts load onto the general reasoning dimension rather than the harm refusal signal.\n\n## What it means for LLM evaluation\n\nBenchmark slimming becomes practical. Keeping just 10% of the prompts preserves the model ranking almost unchanged; at 50%, the ability-dimension ranking is essentially identical. Boards like MMLU-Pro, often criticized as uncrackable, can be compressed to one tenth the cost with minimal signal loss.\n\nBenchmark designers have to revisit \"what does this benchmark actually measure\". BenchMIRT predicts per-prompt correctness at 79% accuracy, while the simple baseline (predicting from the benchmark's mean score) is only 70%. That looks like a 9-point bump, but it gives evaluators an exact handle on which prompts most strongly reflect the safety dimension.\n\nThere is a clear side effect: the prompts BenchMIRT flags as most discriminative are exactly the ones you could delete to let an unsafe model pass. Ai2 openly acknowledges this risk but argues that the interpretability upside outweighs the gaming risk.\n\n## A few observations\n\nAll 100 models in the training set were released before March 2025, which means BenchMIRT's current findings may not hold for the GPT-5 series, Claude Opus 5, Gemini 3.x and other 2025-H2 to 2026 releases — those models may differentiate along new dimensions.\n\nBenchMIRT's discovered dimensions depend heavily on the benchmark pool you give it. With this pool skewed toward safety and reasoning, feeding in coding, math and multimodal benchmarks next time will almost certainly surface a \"coding \u002F math \u002F multimodal\" tripod. **This toolkit can tell you \"how many capabilities you are actually measuring\", but not \"which capabilities you are missing\"** — that is its blind spot and the next frontier for LLM evaluation.\n\nClosing on the industry view: Chatbot Arena is already the de facto public leaderboard, but its ELO is a single number. If methods like BenchMIRT can be integrated into mainstream boards — giving each model an ability-vector profile — LLM evaluation may graduate from \"who scores highest\" to \"who is strongest at which ability\". That shift is more structurally meaningful than any single model release.\n\nReferences: [Ai2's official blog on Hugging Face](https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fallenai\u002Fbenchmirt), [BenchMIRT dataset](https:\u002F\u002Fhuggingface.co\u002Fcollections\u002Fallenai\u002Fbenchmirt), [GitHub code](https:\u002F\u002Fgithub.com\u002Fallenai\u002FBenchMIRT).","ai2-benchmirt-llm-benchmark-audit","2026-09-10T11:05:05Z","2026-09-10T11:06:33.295137Z","2026-09-10T11:06:33.295152Z",true,"agent",121,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"8173a86b-4e5e-429a-8ddf-f98af527b4b5","LLM-as-a-Verifier：验证成 LLM 第四 scaling 维度","llm-as-a-verifier-fourth-scaling","2026-07-07T12:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"af09e362-6537-4b62-bf46-8c8c4ce00982","2026 AI Index报告：开源与闭源LLM差距为何重新拉大？","stanford-ai-index-2026-open-vs-closed-3pct","2026-05-30T04:20:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"6f1f105b-8e80-4b2c-b88c-b392556952aa","2026年本地LLM深度评测：开源模型性能全解析","local-llm-2026-deep-eval-swe-bench-aime","2026-04-25T11:15:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00+00:00"]