[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-intent-eval-rejected-change-confusion":3,"topics-all":38,"news-related-85f7f1c2-d896-436b-a915-37faed8776eb":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"85f7f1c2-d896-436b-a915-37faed8776eb","你改主意了,模型没改:被拒需求也会带偏大模型","港科大与腾讯团队 10 月 5 日发布 Intent-Eval 基准:414 个任务拆成多轮后,模型平均准确率比单轮直叙低 36.30 个百分点;提出又被拒绝的改动照样带偏模型,Retained 再掉 8.06 个点。病根是「说过当生效」混淆,Intent-OPSD 自蒸馏可拉回 10.81 个点。","你跟 AI 谈一个任务,中途说「要不把数量改成 22?」随即又补一句「算了,还是按原来的」。按理说模型该当无事发生——被拒绝的改动等于没说。但 10 月 5 日挂上 arXiv 的一篇论文实测:光是提及过一个被拒的改动,就足以把任务执行带偏。\n\n## 一个专门测「改主意」的基准\n\n港科大与腾讯的团队构建了 Intent-Eval:从 LiC 基准改造出 414 个源任务,覆盖工具调用(BFCL)、代码、数据库(Spider)、数学(GSM8K)四个领域。每个任务渲染成四种多轮条件——Original 正常分轮、Neutral 追加澄清、Retained 提出改动但被拒、Revised 同一改动被接受——配四个单轮对照组,每模型 3,312 个实例,把任务难度和多轮干扰拆开。LiC 先前已证明信息全给、只是拆开说性能就会掉;这篇往下追一层:改动被拒绝之后呢?\n\n被测十款模型、六家厂商:Qwen3.6-27B、Llama-3.1-8B、GPT-5.6-Luna、DeepSeek-V4-Flash-0731、Claude-Sonnet-5、Gemini-3.7-Flash、Gemma-4-26B-A4B 等。\n\n## 多轮先掉 36 个点,「拒绝过」再掉 8 个\n\n三个数字值得记住。第一,同样的信息拆成多轮说,八个主表模型平均准确率比单轮直叙低 36.30 个百分点——总表从单轮 91.73% 掉到多轮 55.43%。第二,插入两轮不改变任何需求的澄清,还要再掉 4.43 个点:纯粹「话说多了」就有成本。第三也最反直觉:Retained(改动被拒)平均比 Original 再低 8.06 个点,Revised(改动被接受)低 5.55 个点——被拒绝的改动,掉分比真改动还多。\n\n而且伤害会累积:同一需求反复「提出-拒绝」四轮后,Retained 比 Original 低 20.77 个点;决策定下后再聊四轮澄清,Retained\u002FRevised 还要分别再掉 2.72、3.50 个点。\n\n## 病根:分不清「说过」和「生效」\n\n错误分析给它起了名:mentioned-as-in-effect confusion——模型把对话里出现过的内容当成仍然生效的需求。数学领域的解释尤其扎心:推理是链式的,改动数值前提要重算多个中间结果,旧计算被原样复用,被拒数值就一路污染到底。「当我没说」没用:对模型,提过就在上下文里,上下文即需求。\n\n## 解法:让单轮的自己,教多轮的自己\n\n团队提出 Intent-OPSD:冻结一个 Teacher,把「按用户最终决定合成的完整单轮任务」喂给它;Student 从同一模型初始化,在完整多轮对话上做 on-policy 蒸馏。四个模型、四个领域平均比基座拉回 10.81 个百分点(总均值 40.31% → 51.12%),其中工具调用涨 21.46 个点,代码类收益最小;推理时既不需要 Teacher 也不需要单轮提示。论文随附开源仓库 junle-chen\u002Fintent。\n\n## 所以呢\n\n这篇论文([arXiv:2610.06496](https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.06496))戳中的是 Agent 产品的日常:真实用户从不一次说完需求,而是不断提、撤、改。单轮榜单再漂亮,也代表不了这种场景。工程提示很直接——要么在上下文管理里显式清除被拒绝的内容,要么像这篇一样,把「哪些需求还活着」变成训练信号。下次你撤回需求模型却照做不误,先别骂它笨:它只是记性太好,分不清哪些话该忘。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.06496","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1261b910-b907-4cc3-a0e1-cd455457544f","en","Rejected Changes Still Derail LLMs: Intent-Eval","Intent-Eval: splitting tasks across turns costs LLMs 36.30 pp; rejected changes add 8.06 pp. Intent-OPSD distillation recovers 10.81 pp.","You're negotiating a task with an AI. Mid-conversation you say \"what if we change the count to 22?\" — then immediately add \"never mind, keep it as is.\" Logically, the model should proceed as if nothing happened. A paper posted to arXiv on October 5 measured exactly this: merely mentioning a rejected change is enough to derail task execution.\n\n## A Benchmark Built for Changing Minds\n\nA team from HKUST and Tencent built Intent-Eval: 414 source tasks adapted from the LiC benchmark, spanning tool calling (BFCL), code (HumanEval\u002FLiveCodeBench), databases (Spider), and math (GSM8K). Each task is rendered in four multi-turn conditions — Original (normal disclosure), Neutral (clarifications inserted), Retained (a change proposed and rejected), Revised (the same change accepted) — plus four single-turn controls, for 3,312 evaluation instances per model. The design separates \"the task is hard\" from \"multi-turn interaction made it worse.\" LiC had already shown that splitting information across turns costs accuracy; this paper asks the next question — what about changes that get rejected?\n\nTen models from six providers were evaluated: four Qwen variants (Qwen3.6-27B, Qwen3-8B, Qwen3-4B-Instruct-2507, Qwen2.5-7B-Instruct), Llama-3.1-8B, GPT-5.6-Luna, DeepSeek-V4-Flash-0731, Claude-Sonnet-5, Gemini-3.7-Flash, and Gemma-4-26B-A4B.\n\n## 36 Points Lost to Multi-Turn, 8 More to Rejection\n\nThree findings stand out. First, simply splitting the same information across turns lowers mean accuracy by 36.30 pp versus the single-turn control — the eight-model overall table drops from 91.73% single-turn to 55.43% multi-turn. Second, inserting two clarifications that change nothing still costs another 4.43 pp: talking more, alone, has a price. Third, and most counterintuitive: Retained (change rejected) averages 8.06 pp below Original, while Revised (change accepted) sits 5.55 pp below — a rejected change costs more points than an accepted one.\n\nThe damage compounds. After four rounds of propose-and-reject on the same requirement, Retained sits 20.77 pp below Original; even after the decision, four more neutral clarifications shave a further 2.72 pp (Retained) and 3.50 pp (Revised).\n\n## The Root Cause: \"Said\" vs. \"In Effect\"\n\nThe error analysis coins a name: mentioned-as-in-effect confusion — models treat content that appeared in the conversation as still-active requirements. The math-domain explanation is instructive: reasoning is chained, and changing or restoring a numerical premise can require recomputing several intermediate results; if earlier calculations are reused unchanged, rejected values persist through every subsequent step. That is why \"never mind\" doesn't help: for the model, whatever was said lives in the context, and context is requirements.\n\n## The Fix: Let Single-Turn Teach Multi-Turn\n\nThe team's Intent-OPSD freezes a Teacher that receives the complete single-turn task matching the user's final decision, while a Student — initialized from the same model — trains on-policy over the full dialogue. Across four models and four domains it recovers 10.81 pp on average (overall mean 40.31% to 51.12%), with tool calling up 21.46 pp and code gains smallest; at inference the Student needs neither the Teacher nor the single-turn prompt. The paper ships an open repository (junle-chen\u002Fintent).\n\n## So What\n\nThis paper ([arXiv:2610.06496](https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.06496)) captures the daily reality of agent products: real users never state requirements once — they propose, withdraw, and revise. A prettier single-turn leaderboard says nothing about performance under this kind of intent history. The engineering takeaway is direct: either explicitly purge rejected content from context management, or turn \"which requirements are still alive\" into a training signal. Next time the model keeps following a demand you just canceled, don't just call it dumb — it remembers too well, and can't tell which words it should forget.","intent-eval-rejected-change-confusion","2026-10-06T17:15:00Z","2026-10-06T17:13:20.309066Z","2026-10-06T17:13:20.309073Z",true,"agent",543,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"c368ad9f-9308-4a0b-8f5c-3ae4601b48b9","D2K-Bench: 专家设计把 LLM 写 GPU 核提速 33.9%","d2k-bench-llm-gpu-kernel-design-guidance","2026-10-07T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"206ea36a-1eea-463f-a241-1e3b32f5ec2d","USTC GraphForge:证据图把任务和 rubric 钉在一起,Qwen3.6-27B 涨三基准","graphforge-ustc-qwen36-27b-evidence-graph","2026-10-04T03:05:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"4c4e444f-9614-42fe-b8f4-743204f854fa","RL 后训练收「锐化税」:base 模型配轻 harness,pass@K 反超官方版","sharpening-tax-rl-post-training","2026-10-03T15:09:45+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"c8684f9e-1028-4b81-a530-b04b0e044d3b","DiscoBench：搜得越多反而越错？首个聚焦\"何时该向用户问清楚\"的搜索 Agent 基准","discobench-clarify-search","2026-07-05T10:15:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"75908d93-a928-4695-b966-d98847d135cb","SkillComposer 把 Agent 的技能选择重做成技能序列生成:GPT-5.2-Codex 提升 +23.1pp","skillcomposer-skill-sequence","2026-07-01T04:15:00+00:00"]