[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-autotool-dynamic-tool-selection":3,"news-related-5083a7bf-ab57-4ddc-900e-096af6d618d0":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"5083a7bf-ab57-4ddc-900e-096af6d618d0","AutoTool 把工具调用做成「动态选择」:训练见 460 工具,推理泛化到 1346 个工具","ICML 2026 入选论文 AutoTool（arXiv:2512.13278）提出「动态工具选择」训练框架，直指现有 Agentic RL「工具集必须固定」的痛点。论文方法分两阶段：Phase I 用 SFT + RL 把「在长 CoT 中插入工具调用」这条轨迹稳定下来；Phase II 用 KL 约束的 Plackett-Luce Ranking 做多步工具选择精修，把「先稳定再精修」这套后训练范式延伸到了工具维度。数据集规模上，他们构建了一个 200K 显式标注的工具调用轨迹数据集，每一步都标注选了哪个工具、为什么选，覆盖 1346 个工具、120 类任务（数学、科学、搜索 QA、代码、多模态都包含）。在 10 个基准上的结果：Qwen3-8B 平均 +6.4%（数学\u002F科学）、+4.5%（搜索 QA）、+7.7%（代码）、+6.9%（多模态），全量开源。最值得展开讲的是「未见工具泛化」实验：训练时模型只暴露 460 个工具，推理时却能在 1346 工具池（含 886 个未见工具）中稳定发挥作用——这把工具调用从「闭集选择」推进到了「开放词汇检索」，对 Agent 生态意义重大。另一条值得注意的是 Qwen2.5-VL-7B 走同一条训练管线也能拿到 +6.9% 多模态增益，说明「该选什么视觉工具」也是可以被学到的。代码、模型、数据全部开源在 GitHub（Gen-Verse\u002FOpen-AgentRL），做 Agent 训练框架的团队都值得通读一遍。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2512.13278","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"caae1fb3-c9bd-43b7-8093-f4805a98d22c","en","AutoTool: trained on 460 tools, generalizes to 1,346","The ICML 2026-accepted paper AutoTool (arXiv:2512.13278) proposes a \"dynamic tool selection\" training framework, directly hitting the pain point of \"the tool set must be fixed\" in current Agentic RL. The method is in two phases: Phase I uses SFT + RL to stabilize the trajectory of \"inserting tool calls in long CoT\"; Phase II uses KL-constrained Plackett-Luce Ranking for multi-step tool-selection refinement, extending the \"stabilize then refine\" post-training paradigm to the tool dimension. In dataset scale, they built a 200K explicitly annotated tool-call trajectory dataset, with each step labeled with which tool was selected and why, covering 1346 tools and 120 task categories (math, science, search QA, code, multimodal all included). Results on 10 benchmarks: Qwen3-8B averages +6.4% (math\u002Fscience), +4.5% (search QA), +7.7% (code), +6.9% (multimodal), fully open-sourced. Most worth elaborating on is the \"unseen-tool generalization\" experiment: at training time the model is exposed to only 460 tools, yet at inference it can stably perform in a 1346-tool pool (including 886 unseen tools) — this pushes tool-calling from \"closed-set selection\" to \"open-vocabulary retrieval\", with major implications for the Agent ecosystem. Another note worth flagging: Qwen2.5-VL-7B walking the same training pipeline also gets +6.9% multimodal gain, showing that \"which visual tool to pick\" is also a learnable thing. Code, model, and data are all open-sourced on GitHub (Gen-Verse\u002FOpen-AgentRL), and any team building an Agent training framework should read it through.","autotool-dynamic-tool-selection","2026-07-12T14:10:00Z","2026-07-12T14:11:25.771513Z","2026-08-19T02:08:40.142862Z",true,"agent",142,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00"]