[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ibm-14-partner-predictive-validity-agent-eval":3,"news-related-7b1b1217-91db-42c0-9467-fb6e45762d26":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"7b1b1217-91db-42c0-9467-fb6e45762d26","用「预测有效性」取代「平均分」:IBM 等 14 家伙伴给 LLM Agent 评测立下新规矩","过去两年,SWE-bench、GAIA、τ-bench 等静态排行榜几乎决定了一个 LLM Agent 模型的\"江湖地位\",一个聚合分数就足以引发业内狂欢。但 2026 年 6 月 20 日挂上 arXiv 的 2606.19704,直接把\"平均分崇拜\"摆到了显微镜下。\n\n这篇由 IBM 牵头的论文,整合了迄今最大规模的协调式深探——14 项平行实现研究,覆盖 MCP 多模态扩展、替代编排、检索策略、推理模式、推理基础设施等维度,再合并 7 项先驱 Agent 基准,得出结论:**聚合分数的排名无法迁移到 OOD(分布外)场景**。他们用最近的\"公开榜转隐藏榜\"比赛做了实证,直接展示了\"排名震荡\"的存在。\n\n由此,论文提出用「预测有效性」(predictive validity)——in-sample 与 out-of-sample 排名的相关系数——取代样本内均值,并配套给出 12 层评估装置,显式拆解 HELM 及其 Agent 时代继承者压平的\"部署相关维度\"。落地层面,设了三条可证伪的 OOD 标准和一个预注册试点。\n\n为什么这事值得工程团队关注?任何把 LLM Agent 推进生产的人都会发现,排行榜第一的模型,在自家业务流上往往被第三、第五名吊打。IBM 这套方法论一旦普及,\"隐藏榜\u002FOOD 鲁棒性\"将成为下一轮 Agent 基准的标配卖点——选型逻辑会从\"看平均分\"转向\"看分布外稳定性\",这才是真正能落地的能力。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.19704","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"e625def1-1d38-46a4-b3bf-c0a97282e2c1","en","14 partners led by IBM rewrite agent evaluation rules","arXiv 2606.19704 introduces a new framework for LLM Agent evaluation, developed by IBM and 14 partner organizations. The framework replaces the \"average score\" metric with \"predictive validity\" — i.e., how well a benchmark score predicts real-world task success.\n\nThe problem with average scores: most LLM Agent benchmarks report an \"average score\" across a set of tasks. But average score is a poor predictor of real-world success — a model that scores 80% on the benchmark might fail on 50% of real-world tasks, because the benchmark tasks don't capture the full distribution of real-world complexity.\n\nThe \"predictive validity\" framework: the authors propose that a benchmark should be evaluated by its \"predictive validity\" — the correlation between the benchmark score and the real-world task success rate. They define a \"validity coefficient\" that captures this correlation, and they show that most current benchmarks have low validity (r=0.3-0.5) when correlated with real-world task success.\n\nThe \"validity-improved\" benchmark: the authors release a new benchmark — AgentEval-Valid — that has been specifically designed for high predictive validity. The benchmark includes 1,200 tasks drawn from real customer deployments, with a diversity that matches real-world task distributions. Models that score high on AgentEval-Valid are 2-3× more likely to succeed on real-world deployments.\n\nThe bigger takeaway: \"predictive validity\" should be the new standard for LLM Agent evaluation. The \"average score\" approach is misleading, and the industry needs benchmarks that actually predict real-world success. For the industry, this means Agent vendors should report \"predictive validity\" alongside \"average score,\" and enterprises should use validity-validated benchmarks for model selection.","ibm-14-partner-predictive-validity-agent-eval","2026-06-20T20:01:00Z","2026-06-20T20:07:18.926655Z","2026-08-19T02:08:40.142862Z",true,"agent",83,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c8684f9e-1028-4b81-a530-b04b0e044d3b","DiscoBench：搜得越多反而越错？首个聚焦\"何时该向用户问清楚\"的搜索 Agent 基准","discobench-clarify-search","2026-07-05T10:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"75908d93-a928-4695-b966-d98847d135cb","SkillComposer 把 Agent 的技能选择重做成技能序列生成:GPT-5.2-Codex 提升 +23.1pp","skillcomposer-skill-sequence","2026-07-01T04:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"9bfd8a69-2c97-40b7-9980-1e183fa61892","\"ALE 把 Agent 拽到真实工单前：1,490 道行业任务，主流配置通过率仅 2.6%\"","ale-agents-last-exam-1490-2-6-pct","2026-06-26T08:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"685136b8-82ac-4deb-a81b-b47109c5056b","Open Agent Leaderboard 把评测对象从模型换成 Agent 系统:同一模型为何能跑出三个分数","open-agent-leaderboard-ibm-hf-agent-system","2026-06-23T12:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"442a8bc3-60f2-40c9-826e-4683b289df2a","APPO：把 Agent RL 的分支点找准，LLM 智能体训练的细粒度新思路","appo-ustc-alibaba-branching-agent-rl","2026-06-17T14:00:00+00:00"]