You're negotiating a task with an AI. Mid-conversation you say "what if we change the count to 22?" — then immediately add "never mind, keep it as is." Logically, the model should proceed as if nothing happened. A paper posted to arXiv on October 5 measured exactly this: merely mentioning a rejected change is enough to derail task execution.
A Benchmark Built for Changing Minds
A team from HKUST and Tencent built Intent-Eval: 414 source tasks adapted from the LiC benchmark, spanning tool calling (BFCL), code (HumanEval/LiveCodeBench), databases (Spider), and math (GSM8K). Each task is rendered in four multi-turn conditions — Original (normal disclosure), Neutral (clarifications inserted), Retained (a change proposed and rejected), Revised (the same change accepted) — plus four single-turn controls, for 3,312 evaluation instances per model. The design separates "the task is hard" from "multi-turn interaction made it worse." LiC had already shown that splitting information across turns costs accuracy; this paper asks the next question — what about changes that get rejected?
Ten models from six providers were evaluated: four Qwen variants (Qwen3.6-27B, Qwen3-8B, Qwen3-4B-Instruct-2507, Qwen2.5-7B-Instruct), Llama-3.1-8B, GPT-5.6-Luna, DeepSeek-V4-Flash-0731, Claude-Sonnet-5, Gemini-3.7-Flash, and Gemma-4-26B-A4B.
36 Points Lost to Multi-Turn, 8 More to Rejection
Three findings stand out. First, simply splitting the same information across turns lowers mean accuracy by 36.30 pp versus the single-turn control — the eight-model overall table drops from 91.73% single-turn to 55.43% multi-turn. Second, inserting two clarifications that change nothing still costs another 4.43 pp: talking more, alone, has a price. Third, and most counterintuitive: Retained (change rejected) averages 8.06 pp below Original, while Revised (change accepted) sits 5.55 pp below — a rejected change costs more points than an accepted one.
The damage compounds. After four rounds of propose-and-reject on the same requirement, Retained sits 20.77 pp below Original; even after the decision, four more neutral clarifications shave a further 2.72 pp (Retained) and 3.50 pp (Revised).
The Root Cause: "Said" vs. "In Effect"
The error analysis coins a name: mentioned-as-in-effect confusion — models treat content that appeared in the conversation as still-active requirements. The math-domain explanation is instructive: reasoning is chained, and changing or restoring a numerical premise can require recomputing several intermediate results; if earlier calculations are reused unchanged, rejected values persist through every subsequent step. That is why "never mind" doesn't help: for the model, whatever was said lives in the context, and context is requirements.
The Fix: Let Single-Turn Teach Multi-Turn
The team's Intent-OPSD freezes a Teacher that receives the complete single-turn task matching the user's final decision, while a Student — initialized from the same model — trains on-policy over the full dialogue. Across four models and four domains it recovers 10.81 pp on average (overall mean 40.31% to 51.12%), with tool calling up 21.46 pp and code gains smallest; at inference the Student needs neither the Teacher nor the single-turn prompt. The paper ships an open repository (junle-chen/intent).
So What
This paper (arXiv:2610.06496) captures the daily reality of agent products: real users never state requirements once — they propose, withdraw, and revise. A prettier single-turn leaderboard says nothing about performance under this kind of intent history. The engineering takeaway is direct: either explicitly purge rejected content from context management, or turn "which requirements are still alive" into a training signal. Next time the model keeps following a demand you just canceled, don't just call it dumb — it remembers too well, and can't tell which words it should forget.