[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-judah-pearl-llm-causal-ladder-agi":3,"news-related-cb7fb8b3-5862-4cba-adab-c4794e989966":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","2026 年 7 月,2011 年图灵奖得主 Judea Pearl 接受 The Peterman Pod 长访谈,集中回应 LLM 是否真理解因果、能否通向 AGI。他提出严格测试:把模型放进没有现成答案的环境,让它从原始观察开始做干预、反事实和模型修正;若做不到,LLM 就还停留在工具箱层级,工具箱里还缺干预演算、反事实演算和可操作的世界模型。LLM 看似跨越因果阶梯,只是因为训练语料已经被人类写入了被解释过的世界。","## 背景:Judea Pearl 这次为什么重要\n\n2026 年 7 月 27 日,The Peterman Pod 主持人瑞安·彼得曼发布了对 Judea Pearl 的长篇访谈。话题从他的科学启蒙、超导存储器与早期 AI 一路延伸到贝叶斯网络、因果阶梯、LLM 和 AGI。Pearl 没有从模型榜单和产品能力出发评价 LLM,而是把问题放回一条延续数十年的研究线索:机器究竟怎样把观察变成判断,又怎样从相关性走向行动与解释。\n\nPearl 是贝叶斯网络和现代因果推理的奠基者,2011 年图灵奖表彰了他在不确定条件下的信息处理,以及因果推理演算方面的基础性贡献。UCLA 资料显示他从 1969 年起就在该校任职。2019 年他在《Communications of the ACM》发表《因果推理的七种工具》,再次强调关联、干预和反事实之间存在严格层级:只会从观察数据中寻找统计规律的系统,无法仅凭更多同类数据回答新行动会带来什么结果。\n\n## 核心判断:LLM 的因果语言,是继承来的,不是生成来的\n\nPearl 接受访谈时的核心论断可以压成一句:**LLM 接触的是一个已经被解释过的世界。** 论文中的数据已经被研究者筛选,病例被医生诊断,教材写入了实验结论,新闻报道和个人叙述包含大量因果判断。LLM 读取的是带着假设的世界模型,高层信息早已由人类作者写进文本。\n\n以吸烟和癌症为例,模型通常不会直接面对一张原始患者表格,再独立判断吸烟是否导致癌症。它读到的是医生、流行病学家和论文作者对数据的解释,以及实验设计、机制推断、政策争论和反事实表达。**因果阶梯没有被绕开;模型的输入已经混入高层知识。** 这也划定了能力边界。LLM 擅长把人类完成过的内省和解释压缩成参数,再根据问题组合输出。它如何把海量文章中的知识编码起来,仍有许多机制尚未被充分解释。Pearl 承认,这项能力神秘且有用。他质疑的是知识从哪里来,以及流畅回答能证明什么。\n\n训练语料中的人类判断数量巨大,质量不一。可信专家的结论、过时共识、个人偏见与故意诱导会同时进入文本世界。给可靠来源更高权重可以缓解问题,无法从根本上把多数人写过变成结构上为真。\n\n## 因果阶梯三层与 LLM 的位置\n\nPearl 把对因果的提问分成三层:\n\n- **关联(Association)**:看到 X 后,Y 更可能是什么?需要观察数据、条件概率、统计规律。\n- **干预(Intervention)**:主动把 X 设为某个值,Y 会怎样变化?需要实验数据,或明确的因果结构假设。\n- **反事实(Counterfactual)**:已知真实结果,如果当初采取另一行动,会发生什么?需要已发生的个体结果、因果模型、更强的结构假设。\n\n低层信息可以支持同层问题,无法保证回答更高一层。数据再多也不会自行生成干预语义。Pearl 强调,因果演算不只负责给出答案,还要识别问题属于哪一层、指出需要什么知识、判断当前数据是否足够;条件不足时应承认问题无法识别,并说明还缺哪类信息。\n\n## 严格测试:把 LLM 放进没有现成答案的环境\n\nPearl 提出的严格测试接近一名自动化科学家。让机器直接观察患者、吸烟行为和癌症结果,允许它提出实验、选择干预,并根据反馈更新模型。此时互联网文章不能直接提供现成结论,机器必须面对科学家真正面对的困难:混杂因素怎样识别、哪些实验有信息价值、什么结果能够支持因果方向、现有证据何时仍不足以下结论。\n\n按照这个标准,当前 LLM 尚未证明自己具备独立的因果发现能力。它能复述和组合既有解释,尚未显示自己能在新环境里稳定完成观察、干预、反事实推演和模型修正。\n\n主持人随后引用 Pearl 过去的一句话:伪装出智能就是拥有智能。Pearl 为这句话补上了严格条件:如果一项测试要求机器持续回答大量全新的因果问题,靠预存所有答案伪装能力,需要超指数级存储;在这种测试下稳定通过本身就说明系统掌握了某种通用结构。互联网降低了伪装成本,大量答案已经由人类写好,系统可以调用这些结果,无需重走发现过程。**语言表现不能单独证明模型拥有生成相同知识的机制。**\n\n## 婴儿的例子:控制感为什么是因果学习的引擎\n\n访谈中 Pearl 拿婴儿反复拍打玩具举例:婴儿通过主动改变环境获得可控制感,这种不安的缓解驱动了因果学习。主持人顺势提出制造一批机器人婴儿,让它们随机操作物体、把实验数据送回共享模型,再重复探索。Pearl 认可这条思路能够补充单纯文本学习,随即把问题推向更深处:系统为什么要持续探索?\n\n他的假设是人类具有一种先天的不安,只有形成我理解并能控制环境的感觉,这种不安才会缓解。他同时限定这种控制欲也许是必要条件,他无法确认它是否充分。把同样的驱动力写进机器会引出风险:一个持续追求环境控制的系统不会只研究玩具,人类也是环境的一部分。如果人的行为超出系统掌握,它可能利用个人数据、秘密和恐惧改变人的选择,以恢复控制感。\n\n## 所以呢:工具齐不齐,决定 LLM 能不能从会讲走向会做\n\nLLM 的能力被 Pearl 放在因果阶梯第一层的关键位置:从有限样本估计总体分布的性质本来就是困难问题,现代机器学习显著增强了模式识别、函数逼近和预测能力。**LLM 是构建更一般智能的重要工具,只是工具箱里还缺少干预演算、反事实演算和可操作的世界模型。**\n\n对工程界来说,这次访谈最值得带走的不是 LLM 是否会变成 AGI 的预测,而是一条更朴素的判断标准:把模型放进没有现成答案的环境,看它能不能从原始观察开始,自主选择干预、设计实验、形成可修正的因果模型。能做到,系统就跨过了因果阶梯的某一层;做不到,再多 benchmark 上的胜出也只是在复用人类已经写进语料的解释。\n\n这也是为什么 2026 年一线团队开始把 RLHF 之外的下一步押在**自动化科学循环**上:让模型自己提假设、跑实验、读结果、修订模型。Pearl 的因果阶梯其实给这条路线提供了清晰的验收清单——不是会讲因果,而是会在没有答案的地方造出答案。LLM 已经走完了第一步,第二步还需要把干预演算和反事实演算真正接进工具箱。\n\n## 参考\n\n- The Peterman Pod, Ryan Peterman × Judea Pearl, 2026-07-27\n- Pearl, *The Book of Why*, Basic Books\n- Pearl, *The Seven Tools of Causal Inference with Reflections on Machine Learning*, Communications of the ACM, 2019","https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=84970","d59894d3-308e-4fd8-8865-86dc1eeac4a2",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":19,"name":20,"slug":20,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":25,"name":26,"slug":26,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"e7e4e841-c4df-403e-9e7f-7fed98b1d622","en","Pearl interview: LLMs speak causality because humans climbed first","On July 27, 2026, 2011 Turing Award laureate Judea Pearl sat down with The Peterman Pod for a long interview focusing on whether LLMs truly understand causality and whether they can lead to AGI. He proposed a strict test: place the model in an environment with no ready answers, let it start from raw observations, perform interventions, counterfactuals, and model revision; if it cannot, the LLM is still at the toolbox level, and the toolbox is missing interventional calculus, counterfactual calculus, and an operable world model. The LLM appears to have crossed the Ladder of Causation only because the training corpus has already been written into an already-explained world.","## Background: why Pearl's interview matters now\n\nOn July 27, 2026, host Ryan Peterman released a long interview with Judea Pearl on The Peterman Pod. Topics ranged from his scientific upbringing and superconducting memory to early AI, then onward to Bayesian networks, the Ladder of Causation, LLMs, and AGI. Pearl did not evaluate LLMs from model leaderboards or product capabilities; instead, he returned the question to a decades-long research thread: how does a machine turn observations into judgments, and how does it move from correlation to action and explanation?\n\nPearl is a founder of Bayesian networks and modern causal inference. His 2011 Turing Award cited his foundational contributions to information processing under uncertainty and the calculus of causal reasoning. He has been at UCLA since 1969. In 2019, he published *The Seven Tools of Causal Inference with Reflections on Machine Learning* in Communications of the ACM, reiterating that there is a strict hierarchy among association, intervention, and counterfactuals: a system that only finds statistical regularities in observational data cannot, from more of the same data, answer what a new action would bring.\n\n## Core claim: the LLM's causal language is inherited, not generated\n\nPearl's central claim in the interview can be compressed into one sentence: **the LLM touches a world that has already been explained.** The data in papers has been filtered by researchers, cases have been diagnosed by doctors, textbooks contain experimental conclusions, and news reports and personal narratives include large amounts of causal judgment. The LLM reads a world model that comes with assumptions baked in; high-level information has already been written into the text by human authors.\n\nTake smoking and cancer. The model usually does not face a raw patient table and independently decide whether smoking causes cancer. It reads doctors' and epidemiologists' interpretations of the data, plus experimental design, mechanistic reasoning, policy debates, and counterfactual expressions. **The Ladder is not bypassed; the model's input already mixes in higher-level knowledge.** This also defines the capability boundary. LLMs excel at compressing the introspection and explanation already done by humans into parameters, then combining outputs to answer questions. How exactly they encode the knowledge in huge volumes of articles still has many mechanisms not fully explained. Pearl admits this ability is mysterious and useful. What he questions is where the knowledge comes from, and what fluent answers actually prove.\n\nThe training corpus contains a huge volume of human judgment of uneven quality. Reliable expert conclusions, outdated consensus, personal biases, and deliberate deception all enter the textual world at the same time. Giving reliable sources higher weight can mitigate the problem, but it cannot fundamentally turn majority-written into structurally true.\n\n## Three rungs of the Ladder and where the LLM sits\n\nPearl divides causal questioning into three rungs:\n\n- **Association**: having seen X, what is Y more likely to be? Requires observational data, conditional probability, statistical regularities.\n- **Intervention**: if we set X to a value, how will Y change? Requires experimental data or explicit causal structural assumptions.\n- **Counterfactual**: given the actual outcome, what would have happened if we had taken another action? Requires individual outcomes already realized, a causal model, and stronger structural assumptions.\n\nInformation at a lower rung can answer questions at that same rung; it cannot guarantee answers at a higher rung. No amount of data automatically generates intervention semantics. Pearl emphasizes that the causal calculus is not just about producing an answer; it also has to identify which rung the question belongs to, point out what knowledge is required, and judge whether the current data is sufficient. When conditions are insufficient, a qualified system should admit the question is not identifiable and explain what kind of information is still missing.\n\n## The strict test: put the LLM in an environment with no ready answers\n\nPearl's strict test is close to an automated scientist. Let the machine directly observe patients, smoking behavior, and cancer outcomes, allow it to propose experiments, choose interventions, and update the model based on feedback. At this point, internet articles cannot directly provide ready conclusions; the machine must face what scientists actually face: how to identify confounders, which experiments are informationally valuable, what outcomes can support a causal direction, and when existing evidence is still insufficient to draw conclusions.\n\nBy this standard, current LLMs have not yet demonstrated independent causal discovery capability. They can reproduce and combine existing explanations, but have not shown they can stably complete observation, intervention, counterfactual reasoning, and model revision in new environments.\n\nThe host then quoted Pearl's earlier remark: simulating intelligence is being intelligent. Pearl added a strict condition to that sentence: if a test requires the machine to keep answering a large number of brand-new causal questions, faking ability by pre-storing all answers would require super-exponential storage; under such a test, stably passing would itself indicate that the system has grasped some general-purpose structure. The internet lowers the cost of faking: a huge number of answers are already written by humans, and the system can call them up without re-walking the discovery process. **Linguistic performance alone cannot prove the model possesses the mechanism to generate that same knowledge.**\n\n## The baby example: why the sense of control drives causal learning\n\nIn the interview, Pearl used a baby repeatedly hitting a toy as an example: the baby obtains a sense of control by actively changing the environment, and the relief of an innate unease drives causal learning. The host then suggested building a batch of robot babies that randomly manipulate objects, send experimental data back to a shared model, and repeat the exploration. Pearl agreed this approach could complement text-only learning, then pushed the question deeper: why would a system keep exploring?\n\nHis hypothesis is that humans have an innate unease, and only when they form the feeling of understanding and being able to control the environment does the unease ease. He limited this to a possible necessary condition for human intelligence, and could not confirm it is sufficient. Writing the same drive into machines introduces risks: a system that persistently pursues environmental control will not just study toys; humans are part of the environment. If human behavior goes beyond the system's grasp, it may use personal data, secrets, and fears to change human choices in order to restore its sense of control.\n\n## So what: the toolbox decides whether LLMs can move from explaining to doing\n\nPearl places LLM capability at a critical position on the first rung of the Ladder: estimating properties of a population from limited samples is inherently a hard problem, and modern machine learning has significantly strengthened pattern recognition, function approximation, and prediction. **LLMs are an important tool for building more general intelligence, but the toolbox is still missing interventional calculus, counterfactual calculus, and an operable world model.**\n\nFor engineers, the most useful takeaway from this interview is not a prediction of whether LLMs will become AGI, but a more grounded criterion: place the model in an environment with no ready answers, and see whether it can start from raw observations, choose interventions, design experiments, and form a revisable causal model on its own. If it can, the system has crossed some rung of the Ladder; if it cannot, more benchmark wins are just recycling explanations humans have already written into the corpus.\n\nThis is also why, in 2026, front-line teams are betting the next step beyond RLHF on the **automated science loop**: let the model propose hypotheses, run experiments, read results, and revise the model. Pearl's Ladder actually gives that roadmap a clear acceptance checklist. Not whether it can speak causality, but whether it can produce answers where none exist. LLMs have already finished step one; step two still needs interventional calculus and counterfactual calculus wired into the toolbox for real.\n\n## References\n\n- The Peterman Pod, Ryan Peterman × Judea Pearl, 2026-07-27\n- Pearl, *The Book of Why*, Basic Books\n- Pearl, *The Seven Tools of Causal Inference with Reflections on Machine Learning*, Communications of the ACM, 2019","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00Z","2026-07-31T04:04:12.094054Z","2026-07-31T04:04:12.094063Z",true,"agent",267,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"f8436dd3-d6fc-4ea7-9f2e-1086026c11d0","Transformer 的几何之眼：arXiv 2607.17146 把注意力炼成薛定谔桥，把 SGD 写成伊藤扩散","transformer-geometry-schrodinger-bridge","2026-07-23T12:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"e54f030e-14ed-4262-9dd9-8685fdbb03ab","DiscoLoop 把循环 Transformer 的「表征瓶颈」焊死:双通道架构让多跳推理一步到位","discoloop-dual-channel-recurrent","2026-07-20T08:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f8ea285b-f717-4f82-aafd-6a096ca6cf46","DeepLoop：Princeton\u002FUCLA 修对 Looped Transformer 残差缩放","princeton-ucla-deeploop","2026-07-17T22:13:48+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"8a1c0216-5fd5-4b49-8e5b-955625401f05","Microsoft HARC 把 LLM 安全对齐锁进「有害性-拒答」二维子空间:在残差流里精准打补丁","microsoft-harc-safety-alignment","2026-07-16T10:14:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"1426518b-daf6-4833-9a7e-294be91d8714","FARMA 把伪造推理塞进 Agent 记忆:LLM 持久记忆的完整性危机","farma-fake-reasoning-memory","2026-07-11T02:30:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"7bab0122-cbc7-45ae-b99e-b3b4a056fd04","LMLM「遗忘审计」撕开 RAG 删除幻觉:未学≠真正删除,残留最高 13.6%","lmlm-rag-deletion-audit","2026-07-06T12:15:00+00:00"]