Agentic prompts are notoriously long and tangled: role instructions, tool schemas, format templates, and the user's request pile into hundreds of tokens. Within that entangled context, how a model decides between "call a tool" and "answer directly" has remained murky. A paper submitted to arXiv on October 7 (2610.09624, accepted as a NeurIPS 2026 Main Conference Poster) offers a remarkably clean answer: the decision can be compressed onto an internal vector, and swapping a single verb flips it.
The Magic of One Verb
The team (Xijie Gong and seven co-authors) distills complex agentic prompts into minimal contrastive pairs: take the same request and replace an execution verb with an analysis verb — write becomes discuss — and the model's tool-call decision reliably flips. Given write, it reaches for the code executor; given discuss, it just talks. That suggests a compact internal state is steering the choice.
They build 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation), then trace the decision to a vector the paper calls μΔ. The vector passes causal tests: it is both necessary (ablate it and decisions break) and sufficient (inject it and decisions follow), and it generalizes beyond the constructed pairs to native multi-turn τ²-Bench trajectories and verb-free requests.
A Tug of War Between Prior and Suppression
The mechanistic decomposition is the most interesting part. Behavioral ablations show that the agent scaffold itself — those role instructions and tool descriptions — establishes a prior favoring tool calls. Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, while execution verbs largely leave it intact. Downstream, scaffold-reading attention heads and MLP features read out the resulting state and turn it into behavior.
Nor is this one model's private quirk. The researchers replicate the same pattern — scaffold sets the prior, verb does the suppression, vector acts as the switch — across seven models from the Qwen, Mistral, and Granite families. The code is open-sourced in the MI4ToolCalling repository.
So What
For agent engineers, the value is moving "will it call the tool" from post-hoc behavioral observation to causal intervention: since μΔ is necessary and sufficient, you can in principle manipulate it at inference time — forcing tool calls or taming over-calling without more prompt superstition. For interpretability research, the paper demonstrates a playbook: when facing entangled agentic contexts, construct minimal contrastive pairs first to obtain a controllable variable, then localize causally, instead of chiseling through raw noise. Next time a model refuses to call a tool for no apparent reason, think of the vector quietly suppressed by a verb — the problem may not be capability, but the switch.
Reference: https://arxiv.org/abs/2610.09624