[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-skill-use-agent-harness-benchmark":3,"news-related-3c6fcf46-f5bb-4136-931c-69cd64216e12":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"3c6fcf46-f5bb-4136-931c-69cd64216e12","Skill-Use 基准揭示 Agent 短板：会做任务，不等于会用 Skill","8 月 5 日提交的一篇论文推出 Skill-Use 基准，用 Trigger、Compliance、Boundary 三个维度测试 Agent 能否在渐进披露下发现、读取并遵守 Skill。基准覆盖 79 个真实 Skill 和 177 个可执行任务；八个模型、两套 Agent harness 的实验显示，领先配置 SU 也只有 0.613，更换 harness 还会改变分数与模型排序。","# Skill-Use 基准揭示 Agent 短板：会做任务，不等于会用 Skill\n\nAgent 生态正在把大量经验封装成 Skill：一个短描述负责提示适用场景，完整文档再规定操作步骤、可用工具和禁止事项。看起来，只要把正确的 Skill 放进工具箱，模型就能照章办事。但 8 月 5 日提交到 arXiv 的论文 **Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?** 给出了更谨慎的答案：模型完成任务的能力，与它主动发现并正确使用 Skill 的能力，并不是一回事。\n\n## 研究把“会用 Skill”拆成三关\n\n论文关注的是渐进披露机制。Agent 一开始只能看到 Skill 的名称、简短描述和文件路径，必须先判断它是否相关，再打开全文并执行流程。研究者因此把能力拆成三个维度：\n\n- **Trigger**：是否主动读取了相关 Skill；\n- **Compliance**：读取后是否遵守规定的步骤；\n- **Boundary**：是否避开文档明确禁止的操作。\n\n三项最终合成 SU 分数，而且只有触发 Skill 后，后续执行才会得到分数。这种设计刻意避开了只看最终答案的评测方式：任务碰巧做对，并不能证明 Agent 真的理解并遵循了操作规程。\n\n## 79 个 Skill、177 个任务：领先配置仍不稳定\n\nSkill-Use 收集了 **79 个真实 Skill**，配套 **177 个可执行任务**，覆盖九个领域。每个任务都放在隔离的 Docker 沙箱中，包含真实文件和完整工具权限；评测记录整条操作轨迹，再用逐项 rubric 检查动作与最终文件。研究团队还用遮蔽 Skill 的方式排除通用要求，并通过多 Agent 对抗审查检查任务范围和评分可验证性。\n\n在八个模型与两套 Agent harness 上，论文报告的领先配置 SU 只有 **0.613**。失败不是单一问题：有的模型能识别 Skill，却执行时偏离流程；有的模型具备执行能力，却根本没有触发 Skill。更关键的是，更换 harness 后，绝对分数和模型排序都会变化。论文据此强调，Skill 使用能力由“模型 + harness”共同决定，不能只归因于底座模型。\n\n## 预加载全文能救触发，但救不了执行\n\n研究把原生渐进披露与“开局直接塞入完整 Skill”做了配对比较。预加载普遍提高 SU，主要原因是 Trigger 上升；但在两种模式都成功触发的样本里，执行质量差距很小。也就是说，全文提前出现，解决的是“模型没想到要打开文档”，不是“模型看完就一定照做”。\n\nSkill 库规模实验也指向同一个入口问题。研究把库扩展到 1、10、20、30 个 Skill，加入干扰项后，主要增加的是“一个也不调用”，而不是选错 Skill。从单个扩到十个时下降最明显，之后变化趋缓。库越大，描述是否清晰、触发条件是否可辨识，就越像真正的系统设计问题。\n\n## 半吊子照流程，可能比不用还差\n\n论文还把触发 Skill 的运行与关闭 Skill 库的同任务基线配对。结果显示，SU 接近 **0.5** 时，Skill 对任务完成的影响才由负转正。低于这一水平，Agent 可能采用了指定工具或格式，却没有把流程执行完整，最后既没满足 Skill，也失去了模型自由解决问题时的完整性。\n\n这给 Agent 工程一个直接提醒：安装更多 Skill 不是免费升级。真正需要优化的是触发描述、过程可观测性、禁区约束，以及 harness 如何展示和调用 Skill。模型排行榜也不能脱离 harness 单独解读。论文全文与实验边界可在 [arXiv 原文](https:\u002F\u002Farxiv.org\u002Fhtml\u002F2608.04828v1) 核对。\n\n**Skill 的价值不在“写进了多少知识”，而在 Agent 能否在正确时机读到它，并把最后一步也做完。**","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2608.04828v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"5eb09533-7400-4efa-ab5c-114043d0f8c6","en","Skill-use benchmark: agents can do tasks but misuse skills","Skill-Use tests whether agents discover, retrieve, and obey skills under progressive disclosure through Trigger, Compliance, and Boundary. Across 79 real skills, 177 executable tasks, eight models, and two agent harnesses, the leading configuration reaches only 0.613 SU, while changing the harness shifts scores and model rankings.","# Skill-Use Exposes an Agent Weakness: Completing Tasks Is Not the Same as Using Skills\n\nAgent ecosystems increasingly package operational knowledge as skills. A short description signals when a skill applies, while the full document specifies the procedure, permitted tools, and forbidden actions. It is tempting to assume that placing the right skill in an agent's library is enough. A paper submitted to arXiv on August 5, **Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?**, argues otherwise: the ability to complete a task and the ability to recognize and correctly apply a skill are distinct.\n\n## Three stages of skill use\n\nThe paper studies progressive disclosure. At the beginning of a task, an agent sees only a skill's name, short description, and file path. It must decide that the skill is relevant, retrieve the full document, and then follow the procedure. Skill-Use therefore separates performance into three dimensions:\n\n- **Trigger** asks whether the agent retrieves the relevant skill.\n- **Compliance** measures whether it follows the prescribed steps.\n- **Boundary** measures whether it avoids explicitly forbidden operations.\n\nThe benchmark combines these dimensions into an SU score. Execution receives credit only after the relevant skill has been triggered. This matters because a correct final answer does not prove that the agent recognized or followed the intended operating procedure.\n\n## 79 skills, 177 tasks, and unstable leading performance\n\nSkill-Use pairs **79 real-world skills** with **177 executable tasks** spanning nine domains. Every task is grounded in real files and runs in an isolated Docker sandbox with full tool access. The benchmark records the entire action trajectory and evaluates it with a multi-item rubric covering actions and final artifacts. The construction process also uses masked-skill screening to remove generic requirements and multi-agent adversarial review to audit task scope and rubric verifiability.\n\nAcross eight models and two agent harnesses, the leading configuration reaches an SU score of only **0.613**. The failures are not all of one kind. Some models recognize a relevant skill but depart from its procedure. Others could comply once a skill is visible but fail to trigger it. The paper also reports that switching the harness changes absolute scores and model rankings. Skill use is therefore a property of the model-harness combination, not simply of the base model.\n\n## Preloading fixes retrieval more than execution\n\nThe researchers compare native progressive disclosure with a setting in which the complete skill document is inserted into the initial context. Preloading raises SU primarily by increasing Trigger. On paired traces where both modes trigger the skill, the execution-quality gap becomes small. In other words, placing the full document in view helps the model notice the skill, but does not materially improve procedural compliance after selection.\n\nA library-size experiment reaches a similar conclusion. The target skill remains available while the library grows from 1 to 10, 20, and 30 entries through randomly sampled distractors. The main increase is in no-skill outcomes rather than wrong-skill selections. The largest visible drop occurs between one and ten skills, with smaller changes afterward. As libraries grow, precise descriptions and distinguishable trigger conditions become core system-design concerns.\n\n## Partial skill use can be worse than no skill\n\nThe paper also pairs triggered skill-enabled runs with baselines on the same tasks and models with the skill library disabled. The effect on task completion turns from negative to positive near an SU score of **0.5**. Below that level, an agent may commit to a prescribed toolchain or format without carrying the procedure through, producing an artifact that satisfies neither the skill nor the unaided task strategy.\n\nThe engineering lesson is direct: installing more skills is not a free capability upgrade. Teams need to improve trigger descriptions, process observability, prohibited-action boundaries, and the way the harness exposes and invokes skills. Model leaderboards should also disclose the harness rather than presenting agent performance as a model-only trait. The full methods and experimental boundaries are available in the [arXiv paper](https:\u002F\u002Farxiv.org\u002Fhtml\u002F2608.04828v1).\n\n**A skill creates value only when an agent finds it at the right moment and carries the procedure through to the final step.**","skill-use-agent-harness-benchmark","2026-08-06T08:00:00Z","2026-08-06T02:19:03.372234Z","2026-08-06T02:19:03.372242Z",true,"agent",149,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"74de194b-9e2c-45ab-aa13-12fe210e66ba","HiGram 给 Agent 记忆加上“路径定位”：先找证据，再改记忆","higram-agent-memory-path-localization","2026-08-05T09:32:43+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"c4ec4625-4a84-4c24-88f0-0ef1beb4f19e","Grok 4.6 发布:61 分追平 GPT-5.6 Sol,把长程 Agent 的 token 账单砍到四分之一","grok-4-6-agentic-cost-frontier","2026-08-14T19:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00"]