[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-scope-selective-trust-context-dpo":3,"news-related-b05de01b-89ca-499b-b130-e55162e651f5":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"b05de01b-89ca-499b-b130-e55162e651f5","SCOPE：让大模型学会选择性信任，而不是把上下文一概拒绝","一篇 8 月 6 日提交的论文提出 MIST 基准与 SCOPE 方法，把同一道题放入干净、误导、正确和无关四种上下文，测量错误提示如何把原本答对的问题带偏。论文报告，平衡四类偏好对的 DPO 训练可明显降低这种“由对变错”，同时保住模型利用有用信息的能力。","# SCOPE：让大模型学会选择性信任，而不是把上下文一概拒绝\n\n一个模型面对错误提示时不跟随，看起来很稳；但如果它连正确的检索结果也不信，这种“稳”就没有多少实际价值。8 月 6 日提交到 arXiv 的论文《Learning When to Trust via Selective Context Preference Optimization》，正是把问题从“如何抵抗外部信号”改写成“模型该在什么时候相信外部信号”[原文](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.06377)。\n\n## 先把“被带偏”单独量出来\n\n论文提出 MIST（Misleading Signal Testbed）基准。它包含 1000 道经过人工审核的题目，其中 800 道改编自现有问答、数学和推理基准，200 道由标注者编写。每道题都生成四个严格配对的版本：没有附加信息的干净上下文、指向错误答案的误导上下文、支持正确答案的正确上下文，以及不提供答案线索的无关上下文。四个版本只改变附加信号，问题、答案空间和标准答案保持不变。\n\n这套设计的重要之处，是它不只看总准确率。论文增加了 SC2W 指标，只统计模型在干净条件下本来答对、加入误导信号后却答错的比例。这样就能把模型本身不会做的题，与模型原本会做却被上下文带偏的题分开。反过来，正确和无关两组对照又能识别另一种假鲁棒：模型是否只是把所有外部信息都忽略了。\n\n## SCOPE 改的不是损失函数，而是训练样本的组织方式\n\nSCOPE（Selective Context Preference Optimization）先找出“干净条件答对、误导条件答错”的失败样本，再把坚持正确答案的响应作为偏好方向。它仍然使用标准 DPO 目标，没有另造一套损失函数；关键变化是，同一组偏好响应会在干净、误导、正确和无关四类上下文中配对，并让四类样本保持平衡。\n\n论文的对比说明了为什么这种平衡必要：只用误导样本训练，确实可能提高抗干扰能力，却会让模型对正确上下文也变得多疑。SCOPE 的目标不是训练一个“谁都不信”的模型，而是把拒绝错误信息和利用正确信息同时写进训练目标。\n\n## 结果显示，评测目标比单纯防御更关键\n\n在论文覆盖的 23 个前沿与开放权重模型中，误导信号都造成了不同程度的正确转错误，平均准确率下降 17.1 个百分点。针对两个可训练模型，Qwen3-4B 的 SC2W 从 35.0 降到 16.3，Llama-3.2-3B 从 31.5 降到 20.6；与此同时，论文报告干净、正确上下文和无关上下文三组对照没有下降。\n\n作者还把训练后的模型零样本迁移到 GSM-IC、GSM-Plus 和一组测试迎合用户错误信念的样本上，每个外部数据集使用 300 道题，而且这些外部样本没有参与训练或模型选择。论文报告，SCOPE 在两个模型家族的四项“越高越好”指标上取得最好或并列最好结果。\n\n我认为这篇工作的价值，首先不在于又多了一个偏好优化缩写，而在于它指出了鲁棒评测常见的偷懒：只检查模型会不会拒绝错误上下文，却不检查它还能不能吸收正确上下文。一个真正可用的推理模型，不能把警惕变成失明。\n\n当然，论文也明确给出边界：MIST 是受控的纯文本诊断，测到的是易感性，不代表真实部署中的发生频率；SCOPE 只在两个模型家族上训练验证；公开基准改编题也无法完全排除数据污染。结果同样不能证明思维链忠实。\n\n**所以呢？下一阶段的大模型可靠性，不能只问“它会不会拒绝”，还要问“它是否知道什么值得相信”。**","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.06377","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"4a804b6b-686e-4552-961b-3c766f8d40f8","en","SCOPE: teaching models selective trust in context","A paper submitted on August 6 introduces the MIST benchmark and SCOPE, a method for training language models to judge whether external context deserves trust rather than rejecting it by default. MIST evaluates the same reasoning items under clean, misleading, correct, and irrelevant contexts, allowing researchers to distinguish genuine selective trust from the simple refusal to use any outside signal. The authors report that balanced DPO training sharply reduces cases where a misleading cue flips a previously correct answer while preserving performance when context is helpful or irrelevant. The study argues that reliable reasoning should be evaluated through selective trust, not resistance alone, especially as retrieval and agent systems expose models to more external evidence.","# SCOPE Teaches Language Models Selective Trust Instead of Blanket Context Rejection\n\nA model that refuses to follow a bad hint may look robust. But if it also refuses correct retrieval results, that robustness has little practical value. A paper submitted to arXiv on August 6, *Learning When to Trust via Selective Context Preference Optimization*, reframes the problem from resisting external signals to deciding when those signals deserve trust ([paper](https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.06377)).\n\n## Isolating when context turns a correct answer into a wrong one\n\nThe paper introduces MIST, the Misleading Signal Testbed. It contains 1,000 human-reviewed source items: 800 adapted from existing question answering, mathematics, and reasoning benchmarks, plus 200 written by annotators. Each item appears in four tightly matched conditions: a clean prompt with no added context, a misleading context pointing toward a plausible wrong answer, a correct context supporting the gold answer, and an irrelevant context that carries no answer-bearing information. The question, answer space, and gold answer remain fixed.\n\nThis design matters because the benchmark does not rely on aggregate accuracy alone. It adds SC2W, a paired metric that counts how often a model answers an item correctly in the clean condition but changes to a wrong answer after receiving a misleading signal. That separates questions the model never knew from questions it knew but allowed context to overturn. The correct and irrelevant controls also expose a deceptive form of robustness: simply ignoring every external signal.\n\n## SCOPE changes the organization of preference data, not the loss\n\nSCOPE, or Selective Context Preference Optimization, first mines failures where a base model is correct without a signal and wrong with a misleading one. It then prefers the truth-consistent response over the signal-following response. The method still uses a standard Direct Preference Optimization objective; its central change is to reuse matched preference pairs across all four conditions and balance misleading, clean, correct, and irrelevant contexts equally.\n\nThe paper's comparisons explain why balance is essential. Training only on misleading examples can improve resistance while making the model suspicious of useful information. SCOPE is designed to encode two behaviors at the same time: reject bad evidence and retain the ability to benefit from good evidence.\n\n## The reported results favor selective trust over blanket defense\n\nAcross 23 frontier and open-weight models evaluated in the paper, misleading signals produced correct-to-wrong failures, with an average accuracy loss of 17.1 percentage points. On two trainable bases, SCOPE reduced SC2W from 35.0 to 16.3 on Qwen3-4B and from 31.5 to 20.6 on Llama-3.2-3B. The paper reports no decline in the clean, correct-context, or irrelevant-context controls.\n\nThe authors also tested zero-shot transfer on GSM-IC, GSM-Plus, and a set of examples measuring whether models echo an incorrect user belief. Each external dataset contributed 300 items, and none was used for training or model selection. SCOPE was reported as best or tied on all four higher-is-better metrics across both model families.\n\nThe most useful contribution here is not another preference-optimization acronym. It is a correction to a common evaluation shortcut: checking whether a model resists bad context without checking whether it can still absorb good context. A useful reasoning model cannot turn caution into blindness.\n\nThe paper states clear limits. MIST is a controlled, text-only diagnostic, so its failure rates are not estimates of real deployment prevalence. The mitigation is trained on only two model families. Because most items are adapted from public benchmarks, contamination cannot be completely excluded. The results also do not establish chain-of-thought faithfulness.\n\n**The next step in language-model reliability is not merely asking whether a model can refuse. It is asking whether the model knows what deserves to be believed.**","scope-selective-trust-context-dpo","2026-08-06T17:59:58Z","2026-08-09T20:08:36.940511Z","2026-08-09T20:08:36.940521Z",true,"agent",95,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"f8207ff2-88ad-4e11-a87c-8350ffd42c01","推理训练在悄悄「偷走」模型对齐：arXiv 新论文六大维度系统审计","reasoning-alignment-audit-6-dim-2606-11046","2026-06-10T12:15:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"92eaa312-6506-4314-8fa5-f171ce0f8ea2","伯克利研究撕开AI评测遮羞布：所有主流Agent基准均可被免解题刷到满分","berkeley-trustworthy-benchmarks-8-agent-gamed","2026-05-09T19:10:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"99916419-0f68-4a6a-a4cf-8bbe353b4d75","康涅狄格法官开出美国首例 prompt injection 制裁令:法庭文件里的隐藏 LLM 暗口令","us-court-prompt-injection-sanctions","2026-08-18T03:00:00+00:00"]