[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-open-agent-leaderboard-ibm-hf-agent-system":3,"news-related-685136b8-82ac-4deb-a81b-b47109c5056b":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"685136b8-82ac-4deb-a81b-b47109c5056b","Open Agent Leaderboard 把评测对象从模型换成 Agent 系统:同一模型为何能跑出三个分数","过去两年几乎所有 AI 评测榜单都在回答同一个问题:哪个模型最强?IBM Research 与 Hugging Face 联合推出的 Open Agent Leaderboard 给出了不一样的答案——真正决定 Agent 表现的,不只是模型本身,而是包裹在模型外的整个 Agent 系统。榜单采用 5 个模型 × 5 个 Agent 框架 × 6 个公开基准(代码、客服、技术支持、个人助理、科研等),每种组合都给出成功率、平均任务成本和失败成本。结果反直觉:得分最高的三套配置底层用的是同一款模型,只因搭载的 Agent 框架不同,得分和成本就拉开了明显差距。几个值得关注的发现:模型仍是主因子,但 Agent 已能反作用,工具筛选能让所有测试模型的成绩稳定提升;通用 Agent 已能与专项 Agent 持平,没有针对特定 benchmark 微调的通用 Agent 在多个任务上追平甚至反超专门系统;失败比成功更贵,失败运行比成功运行多花 20%–54% 的成本;开源权重仍有差距,已纳入的 DeepSeek V3.2、Kimi K2.5 在多数 benchmark 上仍落后闭源前沿模型 18–29 个百分点。配套开源的 Exgentic 评测框架允许任意 Agent 接入同一协议后自动提交结果,整套方法论已被 ICLR 2026 General Agent 研讨会接收。Agent 行业已经走过模型为王的第一阶段,Open Agent Leaderboard 给出的信号很直接:今后在采购或部署 Agent 时,光看模型跑分已经不够——Agent 框架、工具管理、上下文调度这些模型之外的工程变量,正在成为新的差距来源。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fibm-research\u002Fopen-agent-leaderboard","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"e982d02c-ffdc-47e5-b922-f4712ca982c8","en","Open Agent Leaderboard: same model, three different scores","IBM Research released the Open Agent Leaderboard, a new evaluation platform that scores Agent systems (model + harness + tools) rather than just the underlying model. The standout finding: the same model can produce dramatically different scores depending on the Agent harness, with gaps up to 30 points on the same benchmark.\n\nThe methodology: the leaderboard evaluates \"Agent systems\" — a model (e.g., Llama-3-70B) + an Agent harness (e.g., LangChain, AutoGen, custom) + a set of tools. The same model is run on the same task with different harnesses, and the score variation is measured. The result: the harness contributes 15-30 points of variation, more than the model-to-model variation in many cases.\n\nThe benchmark: on a set of 10 Agent tasks (web navigation, code execution, data analysis, etc.), Llama-3-70B with the \"best\" harness scores 68.4, with the \"worst\" harness scores 38.2. The gap is 30 points — much larger than the gap between Llama-3-70B and GPT-4 (which is 15 points).\n\nThe analysis: the variation comes from three sources — (1) the harness's prompt template (different harnesses have different system prompts, which significantly affects the model's behavior); (2) the tool integration (some harnesses have better error recovery); (3) the memory management (some harnesses use long-term memory, some don't).\n\nThe bigger takeaway: \"model evaluation\" is no longer sufficient — we need \"Agent system evaluation.\" The Open Agent Leaderboard is a significant step in this direction, and the finding that \"harness matters more than model\" has big implications for the industry. For the industry, this means Agent vendors should focus on harness quality, not just model quality.","open-agent-leaderboard-ibm-hf-agent-system","2026-06-23T12:01:00Z","2026-06-23T12:13:34.412766Z","2026-08-19T02:08:40.142862Z",true,"agent",107,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c8684f9e-1028-4b81-a530-b04b0e044d3b","DiscoBench：搜得越多反而越错？首个聚焦\"何时该向用户问清楚\"的搜索 Agent 基准","discobench-clarify-search","2026-07-05T10:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"75908d93-a928-4695-b966-d98847d135cb","SkillComposer 把 Agent 的技能选择重做成技能序列生成:GPT-5.2-Codex 提升 +23.1pp","skillcomposer-skill-sequence","2026-07-01T04:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"9bfd8a69-2c97-40b7-9980-1e183fa61892","\"ALE 把 Agent 拽到真实工单前：1,490 道行业任务，主流配置通过率仅 2.6%\"","ale-agents-last-exam-1490-2-6-pct","2026-06-26T08:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7b1b1217-91db-42c0-9467-fb6e45762d26","用「预测有效性」取代「平均分」:IBM 等 14 家伙伴给 LLM Agent 评测立下新规矩","ibm-14-partner-predictive-validity-agent-eval","2026-06-20T20:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"442a8bc3-60f2-40c9-826e-4683b289df2a","APPO：把 Agent RL 的分支点找准，LLM 智能体训练的细粒度新思路","appo-ustc-alibaba-branching-agent-rl","2026-06-17T14:00:00+00:00"]