[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-llm-tool-call-decision-vector":3,"topics-all":41,"news-related-e40de2f9-9ee4-47f3-b1f1-7bf93a4870a9":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"e40de2f9-9ee4-47f3-b1f1-7bf93a4870a9","一个动词翻转工具调用决策,LLM 内部向量现形","NeurIPS 2026 论文把 agent 提示词压成最小对比句对:把 write 换成 discuss,一个动词就稳定翻转模型是否调用工具的决定。研究定位到因果充分且必要的决策向量,机制在 Qwen、Mistral、Granite 七个模型上复现:脚手架建立调用先验,分析动词通过「无需用工具」特征抑制它。","Agent 提示词是出了名的又长又乱:角色指令、工具 schema、格式模板、用户请求混在一起动辄几百 token。在这种高度纠缠的上下文里,模型到底怎么决定\"这轮调工具\"还是\"直接回答\",一直是笔糊涂账。10 月 7 日提交到 arXiv 的一篇论文(编号 2610.09624,已被 NeurIPS 2026 主会接收为 Poster)给出了一个相当干净的答案:这个决策可以被压缩到一个内部向量上,而且换一个动词就能翻转它。\n\n## 一个动词的魔术\n\n研究团队(Xijie Gong 等 8 位作者)的做法是把复杂的 agent 提示词蒸馏成\"最小对比句对\":同一个请求,只把执行类动词换成分析类动词——比如把 write 换成 discuss——模型是否调用工具的决定就会稳定翻转。write 的时候它去调代码执行器,discuss 的时候它直接开聊。这说明动词背后藏着一个紧凑的内部状态在掌舵。\n\n他们围绕 Python、Java、C++ 构建了 500 组这样的配对提示(300 组做机制分析,200 组留作评估),然后把决策追踪到一个向量上,论文里记作 μΔ。这个向量经过了因果层面的检验:既是必要的(消融它,决策就乱),也是充分的(注入它,决策就跟着走),而且不止在构造出的提示对上成立,还泛化到了原生多轮 τ²-Bench 轨迹和不含动词的请求上。\n\n## 先验与抑制的拉锯\n\n更有意思的是机制拆解。行为消融实验显示,agent 脚手架(那些角色指令和工具说明)本身会建立一个\"倾向于调用工具\"的先验;接下来 Transcoder 分解揭示,分析类动词是通过一组\"工具使用没有必要\"的特征把这个先验压下去,而执行类动词基本不碰它。下游还有专门读取脚手架的注意力头和 MLP 特征,负责把压制后的状态读出来变成最终行为。\n\n这套机制不是某个模型的私有小怪癖。研究者在 Qwen、Mistral、Granite 三个家族的七个模型上都复现了同样的模式——脚手架立先验,动词做抑制,向量做开关。论文代码也已开源(仓库 MI4ToolCalling)。\n\n## 所以呢\n\n对做 Agent 工程的人来说,这篇论文的价值在于把\"调不调工具\"从事后行为观察推进到了因果干预层面:既然 μΔ 是充分必要的,理论上就可以在推理时直接操纵它——想强制触发工具调用或者抑制过度调用,不必再苦练提示词玄学。对机制可解释性方向来说,它示范了一条路:面对纠缠的 agent 上下文,先构造最小对比对拿到可控变量,再做因果定位,而不是在原始噪声里硬凿。下次你看到模型无缘无故拒绝调工具,想想那个被动词悄悄压下去的向量——问题可能不在能力,在开关。\n\n参考:https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.09624","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.09624","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":25,"name":26,"slug":26,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"2a9d7e18-e0fc-461f-993b-fd67bff1d3c6","en","One Verb Flips Tool Calls: LLM Decision Vector Found","NeurIPS 2026: swapping one verb flips tool-call decisions in agentic LLMs; the choice traces to a causal vector, replicated across seven models.","Agentic prompts are notoriously long and tangled: role instructions, tool schemas, format templates, and the user's request pile into hundreds of tokens. Within that entangled context, how a model decides between \"call a tool\" and \"answer directly\" has remained murky. A paper submitted to arXiv on October 7 (2610.09624, accepted as a NeurIPS 2026 Main Conference Poster) offers a remarkably clean answer: the decision can be compressed onto an internal vector, and swapping a single verb flips it.\n\n## The Magic of One Verb\n\nThe team (Xijie Gong and seven co-authors) distills complex agentic prompts into minimal contrastive pairs: take the same request and replace an execution verb with an analysis verb — write becomes discuss — and the model's tool-call decision reliably flips. Given write, it reaches for the code executor; given discuss, it just talks. That suggests a compact internal state is steering the choice.\n\nThey build 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation), then trace the decision to a vector the paper calls μΔ. The vector passes causal tests: it is both necessary (ablate it and decisions break) and sufficient (inject it and decisions follow), and it generalizes beyond the constructed pairs to native multi-turn τ²-Bench trajectories and verb-free requests.\n\n## A Tug of War Between Prior and Suppression\n\nThe mechanistic decomposition is the most interesting part. Behavioral ablations show that the agent scaffold itself — those role instructions and tool descriptions — establishes a prior favoring tool calls. Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, while execution verbs largely leave it intact. Downstream, scaffold-reading attention heads and MLP features read out the resulting state and turn it into behavior.\n\nNor is this one model's private quirk. The researchers replicate the same pattern — scaffold sets the prior, verb does the suppression, vector acts as the switch — across seven models from the Qwen, Mistral, and Granite families. The code is open-sourced in the MI4ToolCalling repository.\n\n## So What\n\nFor agent engineers, the value is moving \"will it call the tool\" from post-hoc behavioral observation to causal intervention: since μΔ is necessary and sufficient, you can in principle manipulate it at inference time — forcing tool calls or taming over-calling without more prompt superstition. For interpretability research, the paper demonstrates a playbook: when facing entangled agentic contexts, construct minimal contrastive pairs first to obtain a controllable variable, then localize causally, instead of chiseling through raw noise. Next time a model refuses to call a tool for no apparent reason, think of the vector quietly suppressed by a verb — the problem may not be capability, but the switch.\n\nReference: https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.09624","llm-tool-call-decision-vector","2026-10-09T13:10:00Z","2026-10-09T13:12:41.042417Z","2026-10-09T13:12:41.042435Z",true,"agent",36,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"206ea36a-1eea-463f-a241-1e3b32f5ec2d","USTC GraphForge:证据图把任务和 rubric 钉在一起,Qwen3.6-27B 涨三基准","graphforge-ustc-qwen36-27b-evidence-graph","2026-10-04T03:05:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"44740b4d-8c2c-44fc-8fff-fd89f3fb54ed","12 万美元 token 把 Copilot 运行时从 TypeScript 搬到 Rust","github-copilot-rust-migration-stephen-toub","2026-09-27T11:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"5083a7bf-ab57-4ddc-900e-096af6d618d0","AutoTool 把工具调用做成「动态选择」:训练见 460 工具,推理泛化到 1346 个工具","autotool-dynamic-tool-selection","2026-07-12T14:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"fc533af5-9e8a-43c2-8408-2f931a6aa398","EngramEdit:把事实塞进 LLM 的记忆抽屉","engramedit-decoupled-knowledge-llm","2026-10-09T04:00:00+00:00"]