[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-discobench-clarify-search":3,"news-related-c8684f9e-1028-4b81-a530-b04b0e044d3b":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c8684f9e-1028-4b81-a530-b04b0e044d3b","DiscoBench：搜得越多反而越错？首个聚焦\"何时该向用户问清楚\"的搜索 Agent 基准","当 LLM 搜索 Agent 反复抓取网页却仍给错答案时,病灶往往不在搜索,而在于它拒绝在用户查询模糊时主动澄清。腾讯 Hunyuan 与清华大学在 arXiv:2606.27669 联合推出 DiscoBench,用 211 个样本、463 处歧义实例和一台\"用户模拟器\",系统揭示了主流大模型在多轮深搜场景下的失败模式。实验覆盖 Gemini-3.1-Pro、Doubao-Seed-2.0-Pro、DeepSeek-V4-Pro、Claude-Opus-4.7 等主流系统,核心结论是主动澄清比反复检索更有效,而搜得更多甚至不如直接猜。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.27669","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d1a47607-c762-474e-a270-8e8f29fc555c","en","DiscoBench: when should search agents ask the user?","When LLM search Agents repeatedly scrape web pages but still give wrong answers, the problem often isn't search, but that they refuse to proactively clarify when the user query is ambiguous. Tencent Hunyuan and Tsinghua University jointly release DiscoBench in arXiv:2606.27669, using 211 samples, 463 ambiguity instances, and a \"user simulator\" to systematically reveal the failure modes of mainstream large models in multi-turn deep-search scenarios. Experiments cover mainstream systems like Gemini-3.1-Pro, Doubao-Seed-2.0-Pro, DeepSeek-V4-Pro, Claude-Opus-4.7. The core conclusion is that proactive clarification is more effective than repeated retrieval, and searching more is even worse than just guessing directly.","discobench-clarify-search","2026-07-05T10:15:00Z","2026-07-05T10:16:08.409030Z","2026-08-19T02:08:40.142862Z",true,"agent",82,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"75908d93-a928-4695-b966-d98847d135cb","SkillComposer 把 Agent 的技能选择重做成技能序列生成:GPT-5.2-Codex 提升 +23.1pp","skillcomposer-skill-sequence","2026-07-01T04:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"9bfd8a69-2c97-40b7-9980-1e183fa61892","\"ALE 把 Agent 拽到真实工单前：1,490 道行业任务，主流配置通过率仅 2.6%\"","ale-agents-last-exam-1490-2-6-pct","2026-06-26T08:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"685136b8-82ac-4deb-a81b-b47109c5056b","Open Agent Leaderboard 把评测对象从模型换成 Agent 系统:同一模型为何能跑出三个分数","open-agent-leaderboard-ibm-hf-agent-system","2026-06-23T12:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7b1b1217-91db-42c0-9467-fb6e45762d26","用「预测有效性」取代「平均分」:IBM 等 14 家伙伴给 LLM Agent 评测立下新规矩","ibm-14-partner-predictive-validity-agent-eval","2026-06-20T20:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"442a8bc3-60f2-40c9-826e-4683b289df2a","APPO：把 Agent RL 的分支点找准，LLM 智能体训练的细粒度新思路","appo-ustc-alibaba-branching-agent-rl","2026-06-17T14:00:00+00:00"]