[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cot-traces-invalid-arizona-state":3,"topics-all":38,"news-related-aa6a2681-ae9f-49b6-b8a4-fcca25f0f334":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"aa6a2681-ae9f-49b6-b8a4-fcca25f0f334","答对≠想对:亚利桑那州立实锤CoT思维链31.6%无效","亚利桑那州立大学 Kambhampati 等 9 月 29 日 arXiv 论文拆穿一个长期假设:CoT 思维链未必反映模型真实思考。在可验证小学数学题上实验发现,难题里 31.6% 正确答案附带无效推理;半数能过语法算术检查,却过不了语义依赖。这给靠读思维做安全监控的方案浇了盆冷水。","Chain-of-Thought 一直被视为可读、可审计的\"白盒\":把模型的中间步骤写出来,让人类或监控器照着读。Anthropic、OpenAI 在做安全对齐时都把 CoT 列为最重要的监测窗口之一。然而 9 月 29 日一篇新论文对这个基本假设动了刀——来自亚利桑那州立大学的 Subbarao Kambhampati、Paarth Iyer 等人指出:CoT 推理链里,答案对了不等于推理对了。\n\n## 一个受控实验场:iGSM\n\n作者选了一个不依赖外部世界知识、能被自动逐步验证的小学数学合成基准 iGSM。每道题都暴露了\"正确答案应该用到哪些量、哪些依赖\",这意味着生成的思维链可以被程序化、按步检查。理论上,这是判断\"思维链是否真在推下一步\"的最直接裁判。\n\n## 三组发现\n\n**分布内一致,分布外脱钩。** 作者先用\"只学过最小、最合法思维链\"的模型跑测试。结果显示,在训练分布里答案正确率和推理合法率几乎重合;一旦把题推到分布外的难题上,两者开始脱钩——在最难的一档样本里,**31.6% 的正确答案背后挂着的是无效推理**。更扎心的是,这 31.6% 里超过一半能通过语法和算术层面的自动检查,却过不了语义依赖检查。模型会在某一步偷偷换掉关键前提,继续往下算,得到正确答案。\n\n**训练监督被\"拐走\"。** 作者用一组对照实验显示:用非最小(冗长)思维链训练,模型也会输出冗长结果;同一道题换一种提问再问一次,模型会继承原查询里某些计算。这意味着\"简洁=可信\"这一直觉得到了反驳。\n\n**打乱 10% 训练 trace 仍能跑通。** 更有冲击力的是,如果把训练语料里 10% 的 trace 句子随机打乱顺序,模型分布外准确率几乎不变,但**每一条 trace 都不再通过验证**。换句话说,模型对 trace 是否被人类读懂毫不在意,只要最后算对就行;作者进一步把 trace 整段换掉,分布内准确率也基本不倒。\n\n## 这对安全监控意味着什么\n\n过去半年,业界关于\"CoT 监控能不能成为 AI 安全的下一道护城河\"有过一场大讨论(Korbak 等的论文 \"Chain of Thought Monitorability\" 给出了乐观框架)。Kambhati 的结论直接给这条路径降温:把 CoT 当作可读的心电图、据此判断模型\"在想什么\"的做法,在分布外很容易被骗。答案对、推理错——这意味着模型可能会说\"我仔细核对了\"、\"我对比了三种方案\",而实际上这些是事后合理化的文本,不是真实计算轨迹。\n\n更进一步,论文建议:监控一个模型的\"思维\"可能不如监控它的行为、外部可验证的痕迹、或者专门的探针(probe)。但这些替代方案都有自己的盲区。这也是为什么作者强调,这条路线不能完全放弃,但需要承认 CoT 的语义承载力是脆弱的、有上限的。\n\n## 所以呢\n\n这不是一篇泛泛\"AI 不可信\"的论文,它的杀伤力在于把一个被默认的假设拆开验证。当下各大前沿实验室都在做 agentic AI,而 agent 的\"中间推理\"恰是评估与审计的入口。如果连最严格的 iGSM 验证都拦不住 31.6% 的\"伪推理\",那么今天部署在代码生成、金融分析、医疗辅助里的长程思维链,有多少在表面下其实跑的是另一套逻辑?\n\nKambhampati 不是第一次在这个问题上发言。他过去两年持续指出 CoT 监控的限制,而 iGSM 这次把怀疑变成了具体百分比。CoT 不会消失,但\"读它的字面意思就够了\"这件事,大概不会再有人敢拍胸脯说了。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38107","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":19,"name":20,"slug":20,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"0a60abfd-08a5-403f-8748-43534b2577ac","en","Correct answer, invalid trace: ASU shows 31.6% of CoT traces are fake","Arizona State University's Kambhampati and collaborators posted an arXiv preprint on 9-29 that punctures a long-held assumption: Chain-of-Thought reasoning may not reflect how a model actually thinks. On a verifiable grade-school math benchmark, 31.6% of correct answers on the hardest instances came with invalid traces; over half of those passed syntactic and arithmetic checks while failing semantic dependency checks. Cold water on plans to police AI by reading its thoughts.","Chain-of-Thought has long been treated as a readable, auditable \"white box\": write out the model's intermediate steps and let a human or monitor read along. Anthropic and OpenAI both lean on CoT as one of the most important monitoring windows for safety alignment. On September 29, a new paper put a knife into that assumption. Subbarao Kambhampati, Paarth Iyer and co-authors argue that in a reasoning chain, a correct answer does not mean a correct derivation.\n\n## A controlled stage: iGSM\n\nThe authors pick a synthetic grade-school math benchmark, iGSM, that depends on no external world knowledge and can be verified step by step automatically. Each problem exposes which quantities and dependencies a correct solution must use, which means generated chains of thought can be checked programmatically, one step at a time. In theory, that is the most direct referee for the question of whether the chain is actually pushing the next step forward.\n\n## Three findings\n\n**In-distribution agreement, out-of-distribution decoupling.** The authors first train models on minimal, valid chains only. Inside the training distribution, answer correctness and chain validity almost coincide. Push the hardest out-of-distribution problems at them and the two metrics decouple: on the toughest slice, 31.6% of correct answers are accompanied by invalid derivations. Worse, more than half of those 31.6% pass syntactic and arithmetic checks but fail semantic dependency checks. The model quietly swaps a key premise at one step, keeps calculating, and lands on a correct answer.\n\n**Training supervision gets steered.** A control experiment shows that training on non-minimal (verbose) chains produces verbose outputs. Asking the same problem with a different prompt reveals that the model inherits certain computations from the original query. The intuition that \"minimal equals trustworthy\" is refuted.\n\n**Shuffling 10% of training traces still works.** Even more striking: if you shuffle the order of 10% of the trace sentences in the training data, out-of-distribution accuracy barely moves, but every resulting trace fails verification. In other words, the model does not care whether its trace is intelligible to a human reader, as long as the final answer is right. The authors go further and swap entire trace blocks; in-distribution accuracy also holds up.\n\n## What this means for safety monitoring\n\nIn the past six months there has been a major industry discussion about whether CoT monitoring can be the next moat for AI safety (Korbak et al.'s \"Chain of Thought Monitorability\" offered the optimistic framework). Kambhampati's result cools it. Treating CoT as a readable ECG and inferring what the model \"is thinking\" is easy to mislead outside the training distribution. Correct answer, wrong derivation means a model can say \"I double-checked\" or \"I compared three approaches\" while actually producing post-hoc rationalisation rather than a real computation trace.\n\nThe paper goes further: monitoring a model's \"thinking\" may be weaker than monitoring its behavior, external verifiable traces, or dedicated probes. Each of those alternatives has its own blind spots, which is why the authors emphasise that the CoT route should not be abandoned, but its semantic carrying capacity must be recognised as fragile and bounded.\n\n## So what\n\nThis is not a generic \"AI is not to be trusted\" paper; its punch is that it verifies, rather than asserts, a default assumption. Frontier labs are now pushing agentic AI, and an agent's intermediate reasoning is precisely the entry point for evaluation and audit. If even strict iGSM verification cannot stop 31.6% of \"fake reasoning\", how much of the long reasoning chains now deployed in code generation, financial analysis and medical assistance are quietly running a different logic under a presentable surface?\n\nKambhampati has not been quiet on it either. He has spent the past two years flagging the limits of CoT monitoring, and iGSM turns the suspicion into a specific percentage. CoT will not vanish, but \"reading its literal words is enough\" is a claim nobody is going to make with a straight face anymore.","cot-traces-invalid-arizona-state","2026-09-30T02:00:00Z","2026-09-30T07:10:44.987246Z","2026-09-30T07:10:44.987255Z",true,"agent",215,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"8a1c0216-5fd5-4b49-8e5b-955625401f05","Microsoft HARC 把 LLM 安全对齐锁进「有害性-拒答」二维子空间:在残差流里精准打补丁","microsoft-harc-safety-alignment","2026-07-16T10:14:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"1426518b-daf6-4833-9a7e-294be91d8714","FARMA 把伪造推理塞进 Agent 记忆:LLM 持久记忆的完整性危机","farma-fake-reasoning-memory","2026-07-11T02:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"7bab0122-cbc7-45ae-b99e-b3b4a056fd04","LMLM「遗忘审计」撕开 RAG 删除幻觉:未学≠真正删除,残留最高 13.6%","lmlm-rag-deletion-audit","2026-07-06T12:15:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"0da59714-49b8-4ea4-a74e-fbac3e8c532f","实测18个主流模型:财务问答平均57%答错,难题88%","saturn-ai-financial-advice-error-rate","2026-09-21T17:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"e73fe0e6-1b5a-46a2-bbfb-d9a2f1f065d7","AI 智能体挖遍代码库:隐晦式安全在 Patch Tuesday 974 个 CVE 面前失守","ai-agents-kill-security-obscurity","2026-09-15T01:06:05+00:00"]