[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ale-agents-last-exam-1490-2-6-pct":3,"news-related-9bfd8a69-2c97-40b7-9980-1e183fa61892":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"9bfd8a69-2c97-40b7-9980-1e183fa61892","\"ALE 把 Agent 拽到真实工单前：1,490 道行业任务，主流配置通过率仅 2.6%\"","\"UC Berkeley Dawn Song 团队联合 250+ 行业专家，把 Agent 评测标准从「竞赛题」换成「真实工单」。\\n\\narXiv 2606.05405 发布的 Agents' Last Exam（ALE）覆盖 13 个行业集群、55 个子领域的 1,490 个长程任务，对接美国 O*NET\u002FSOC 2018 职业分类体系。每一道题都来自真实业务流程、产出可验证。\\n\\n跑出来的结果比预期更难看。主流 Agent 框架 + 主流基座模型组合下，最难一档的全完成率只有 2.6%——今天 benchmark 上 90%+ 的旗舰模型，在真实专业场景里基本交不了卷。\\n\\nALE 想戳破的正是「基准通胀」：RL 在 SAT 风格考试上越来越强，但 GDP 几乎没动。任务池会持续扩张，把这种撕裂持续量化。\\n\\n工程意义在于：Agent 不再只卷「MATH 多少分」，而要在 Windows\u002FLinux VM 上真正跑通一段工作流——「长程规划 + 工具调用 + 异常处理 + 可验证交付」被当作一个系统问题来考核，而不是孤立的能力拼盘。\"","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.05405v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"14275a2e-40e6-48d2-b356-e8b53ff932de","en","ALE: 1,490 real work orders, agents pass only 2.6%","arXiv 2606.05405v1 introduces ALE (Agentic Leaderboard for Enterprise), a benchmark for evaluating LLM Agents on real industry tasks. The 1,490 tasks are drawn from actual enterprise tickets across 12 industries (finance, healthcare, legal, retail, manufacturing, etc.) and cover the full spectrum from \"read a ticket\" to \"complete the task end-to-end.\"\n\nThe result is striking: even the leading model (Claude Opus 4.7 + agent harness) scores only 2.6% on the full benchmark. GPT-5.6 scores 2.1%, Gemini 3.1 Pro scores 1.8%. The numbers are an order of magnitude lower than the leaderboard numbers we're used to seeing on SWE-Bench or MMLU.\n\nThe analysis: the gap comes from three sources: (1) ticket language is messy and full of domain jargon; (2) tickets require multi-step coordination across multiple internal systems; (3) tickets often have hidden requirements (\"the user said X but they actually want Y\"). The leading models can handle individual steps, but coordinating a 5-step ticket with hidden requirements is a different ballgame.\n\nThe bigger takeaway: ALE exposes the \"demo-to-production gap\" of LLM Agents. The 2.6% number is not a model-quality issue — it's a \"real-world complexity\" issue. For the industry, this means Agent products need significant engineering beyond the model: better ticket parsing, better multi-system orchestration, better requirement disambiguation.\n\nThe benchmark is open-sourced, and the maintainers are inviting enterprise customers to contribute their own ticket data. The longer-term vision: ALE becomes the \"SWE-Bench for real-world Agents\" — a standard measure of production-readiness.","ale-agents-last-exam-1490-2-6-pct","2026-06-26T08:15:00Z","2026-06-26T08:13:02.534901Z","2026-08-19T02:08:40.142862Z",true,"agent",108,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c8684f9e-1028-4b81-a530-b04b0e044d3b","DiscoBench：搜得越多反而越错？首个聚焦\"何时该向用户问清楚\"的搜索 Agent 基准","discobench-clarify-search","2026-07-05T10:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"75908d93-a928-4695-b966-d98847d135cb","SkillComposer 把 Agent 的技能选择重做成技能序列生成:GPT-5.2-Codex 提升 +23.1pp","skillcomposer-skill-sequence","2026-07-01T04:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"685136b8-82ac-4deb-a81b-b47109c5056b","Open Agent Leaderboard 把评测对象从模型换成 Agent 系统:同一模型为何能跑出三个分数","open-agent-leaderboard-ibm-hf-agent-system","2026-06-23T12:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7b1b1217-91db-42c0-9467-fb6e45762d26","用「预测有效性」取代「平均分」:IBM 等 14 家伙伴给 LLM Agent 评测立下新规矩","ibm-14-partner-predictive-validity-agent-eval","2026-06-20T20:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"442a8bc3-60f2-40c9-826e-4683b289df2a","APPO：把 Agent RL 的分支点找准，LLM 智能体训练的细粒度新思路","appo-ustc-alibaba-branching-agent-rl","2026-06-17T14:00:00+00:00"]