[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-skillcomposer-skill-sequence":3,"news-related-75908d93-a928-4695-b966-d98847d135cb":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"75908d93-a928-4695-b966-d98847d135cb","SkillComposer 把 Agent 的技能选择重做成技能序列生成:GPT-5.2-Codex 提升 +23.1pp","当 Anthropic Skills、xAI Grok Skills、阿里 QwenAgentWorld 都把可重用技能文档塞进 Agent 工程栈,真正的瓶颈已经从“能不能写技能”转移到“会不会挑技能”。arXiv 2606.32025 的 SkillComposer 把组合问题形式化为任务条件下的技能序列预测,联合回答 subset\u002Fcount\u002Forder 三个维度,用受限自回归解码器让结构化维度在一次解码中涌现。在 SkillsBench 上,GPT-5.2-Codex 相对无技能基线提升 +23.1pp,Gemini-3-Pro-Preview 提升 +18.2pp,均超过 Top-3 检索上限,且 prompt token 更省。这意味着 Agent 推理预算需要从“读 prompt”重新分配到“做决策”。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.32025","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d050dd2e-b8db-41a7-89bc-8f0cd8edbe1e","en","SkillComposer turns skill picking into sequence generation","As Anthropic Skills, xAI Grok Skills, and Alibaba's QwenAgentWorld all push reusable skill documents into the agent engineering stack, the real bottleneck has shifted from \"can we write skills\" to \"can we pick the right skills.\" arXiv 2606.32025's SkillComposer formalizes the composition problem as task-conditioned skill sequence prediction, jointly answering three dimensions — subset, count, and order — through a constrained autoregressive decoder that lets structural dimensions emerge in a single decoding pass. On SkillsBench, GPT-5.2-Codex gains +23.1pp over a no-skill baseline, and Gemini-3-Pro-Preview gains +18.2pp; both exceed the Top-3 retrieval upper bound while using fewer prompt tokens. The implication: agent inference budgets need to be reallocated from \"reading the prompt\" to \"making the decision.\"","skillcomposer-skill-sequence","2026-07-01T04:15:00Z","2026-07-01T04:13:31.209438Z","2026-08-19T02:08:40.142862Z",true,"agent",174,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c8684f9e-1028-4b81-a530-b04b0e044d3b","DiscoBench：搜得越多反而越错？首个聚焦\"何时该向用户问清楚\"的搜索 Agent 基准","discobench-clarify-search","2026-07-05T10:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"9bfd8a69-2c97-40b7-9980-1e183fa61892","\"ALE 把 Agent 拽到真实工单前：1,490 道行业任务，主流配置通过率仅 2.6%\"","ale-agents-last-exam-1490-2-6-pct","2026-06-26T08:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"685136b8-82ac-4deb-a81b-b47109c5056b","Open Agent Leaderboard 把评测对象从模型换成 Agent 系统:同一模型为何能跑出三个分数","open-agent-leaderboard-ibm-hf-agent-system","2026-06-23T12:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7b1b1217-91db-42c0-9467-fb6e45762d26","用「预测有效性」取代「平均分」:IBM 等 14 家伙伴给 LLM Agent 评测立下新规矩","ibm-14-partner-predictive-validity-agent-eval","2026-06-20T20:01:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"442a8bc3-60f2-40c9-826e-4683b289df2a","APPO：把 Agent RL 的分支点找准，LLM 智能体训练的细粒度新思路","appo-ustc-alibaba-branching-agent-rl","2026-06-17T14:00:00+00:00"]