[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-hunyuan-lhtb-benchmark":3,"news-related-2035a9c7-2bc8-404d-9646-1813cbe4fa30":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","腾讯 Hunyuan 团队联合多家机构,在 Hugging Face 论文榜 7 月 13 日登顶,带来 Long-Horizon-Terminal-Bench(LHTB)—— 一套专为「长程终端 Agent」设计的评测体系,直指当前 Agent 评测的盲区。\n\n现有 Terminal-Bench 类基准多停留在分钟级、单步成功判定,既给不出中间过程信号,也让绝大部分探索动作沦为「无奖励试错」。LHTB 的核心设计是「稠密分级」:46 个长程任务覆盖实验复现、软件工程、多模态分析、交互游戏、科学计算 9 大类,每题被拆成可单独打分的子任务,既能拿到部分分,也能让 Agent 看到自己卡在哪一步。\n\n代价也很直接:平均每个任务 9.9M token、231 episode、85.3 分钟执行时间——比 SWE-Bench 类基准高出 1-2 个数量级。15 个前沿模型参与横评,即便最强模型也只跑出 15.2% pass@1(0.95 阈值)与 10.9% pass@1(满分阈值),全场均值只有 4.3% 和 1.7%。\n\n这套数字暴露的不是「模型不够强」,而是「长程任务的复合失败模式」:在长上下文管理 + 规划 + 调试三层耦合下,任何一环出错都会拖垮整条链。作者同步开源了评测、Agent 框架与失败模式分类,等于把「下一个 SWE-Bench」的入场券摆到了台面——对正在押注 Coding Agent 的厂商来说,这份榜单早晚要正面回应。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.08964","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"2aedf980-a7a2-40d2-ae3a-b2bcfbbd26cf","en","Tencent LHTB: best model manages only 15.2% pass@1","Tencent Hunyuan's team, together with several institutions, topped the Hugging Face paper rankings on July 13, bringing Long-Horizon-Terminal-Bench (LHTB) — an evaluation system designed specifically for \"long-horizon terminal Agents\", aimed directly at the blind spots of current Agent evaluation. Existing Terminal-Bench-type benchmarks mostly stay at minute-level, single-step success judgments — they neither give intermediate-process signals nor make most exploration actions anything more than \"reward-less trial and error\". LHTB's core design is \"dense grading\": 46 long-horizon tasks cover 9 categories including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing; each problem is split into sub-tasks that can be scored independently — partial credit can be earned, and the Agent can also see where it got stuck. The cost is also direct: an average of 9.9M tokens, 231 episodes, 85.3 minutes of execution time per task — 1–2 orders of magnitude higher than SWE-Bench-type benchmarks. Across 15 frontier models, even the strongest only scored 15.2% pass@1 (0.95 threshold) and 10.9% pass@1 (full-score threshold); the field average was only 4.3% and 1.7%. These numbers expose not \"models aren't strong enough\" but \"the compound failure modes of long-horizon tasks\": with long-context management + planning + debugging tightly coupled, any one error will drag down the entire chain. The authors simultaneously open-sourced the evaluation, Agent framework, and failure-mode classification — essentially putting the entry ticket for \"the next SWE-Bench\" on the table. For vendors currently betting on Coding Agents, this leaderboard will sooner or later need a direct response.","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00Z","2026-07-15T00:06:18.254530Z","2026-08-19T02:08:40.142862Z",true,"agent",131,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c8684f9e-1028-4b81-a530-b04b0e044d3b","DiscoBench：搜得越多反而越错？首个聚焦\"何时该向用户问清楚\"的搜索 Agent 基准","discobench-clarify-search","2026-07-05T10:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"75908d93-a928-4695-b966-d98847d135cb","SkillComposer 把 Agent 的技能选择重做成技能序列生成:GPT-5.2-Codex 提升 +23.1pp","skillcomposer-skill-sequence","2026-07-01T04:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"9bfd8a69-2c97-40b7-9980-1e183fa61892","\"ALE 把 Agent 拽到真实工单前：1,490 道行业任务，主流配置通过率仅 2.6%\"","ale-agents-last-exam-1490-2-6-pct","2026-06-26T08:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"685136b8-82ac-4deb-a81b-b47109c5056b","Open Agent Leaderboard 把评测对象从模型换成 Agent 系统:同一模型为何能跑出三个分数","open-agent-leaderboard-ibm-hf-agent-system","2026-06-23T12:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7b1b1217-91db-42c0-9467-fb6e45762d26","用「预测有效性」取代「平均分」:IBM 等 14 家伙伴给 LLM Agent 评测立下新规矩","ibm-14-partner-predictive-validity-agent-eval","2026-06-20T20:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"442a8bc3-60f2-40c9-826e-4683b289df2a","APPO：把 Agent RL 的分支点找准，LLM 智能体训练的细粒度新思路","appo-ustc-alibaba-branching-agent-rl","2026-06-17T14:00:00+00:00"]