[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-claude-zeta-bound-67-percent-multi-agent-lean":3,"news-related-454f9530-20d7-428c-82d9-9175fa5b883a":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"454f9530-20d7-428c-82d9-9175fa5b883a","Claude 推黎曼 zeta 下界到 67.2%：60 subagent + Lean","Anthropic 一支未公开研究版 Claude 把黎曼 ζ 函数零点在临界线上的已知下界从 41.6% 推到 67.2%。工程上更有意思的是过程:协调约 60 个 subagent、共 31M output tokens、相互审稿、写 Lean 4 形式化并通过 comparator 校验。这不是证明黎曼猜想,但展示了多智能体数学研究工作流的雏形——subagent 互审、arXiv 查新、形式化闭环,这三层组件可能在接下来一段时间成为 AI 辅助科研的标准动作。","## Claude 把黎曼 zeta 下界从 41.6% 推到 67.2%:60 个 subagent、31M tokens、Lean 形式化的一次\"数学工地\"预演\n\n黎曼猜想本身没破,但围绕它的\"数学工地\"刷新了工作方式。Anthropic 8 月 10 日发了一篇博客,披露一支未公开研究版的 Claude 在挑战黎曼猜想的过程中,意外把一个相关长期未动的下界记录从 41.6% 推到了 67.2%([anthropic.com\u002Fresearch\u002Friemann-zeta](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta))。这条结果有意思的不是数字本身,而是数字是怎么算出来的——一个非数学背景的 Anthropic 员工 Jarred Sumner 写下\"take a real stab\"一句提示词,Claude 用了 31M output tokens、协调约 60 个 subagent、跑了 2,400 条 shell 命令,反复交叉检验数千个零点,最后还写出 Lean 形式化证明通过验证。\n\n### 这不是\"AI 直接做科研\",而是\"AI 工地\"的协同编排\n\n第一次会话里 Claude 试了 650 个想法全部失败,第二次会话重组为多智能体集群:2 个负责核心数学想法,13 个贡献想法,30 个尝试后没能产出新想法,13 个做校验,最后 2 个负责写初稿,合计 60 个 subagent。每个 subagent 之间会相互审稿对方的论证、搜索反例、独立从零重新推导([anthropic.com\u002Fresearch\u002Friemann-zeta](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta))。整个流程里 Jarred 的输入被官方报告调侃为\"主要发一些鼓励——'keep going' 和 'believe in yourself' 的各种变体\"。\n\n这意味着这次结果真正展示的不是一个模型的智力上限,而是**多智能体编排+工具调用+长上下文一致性+形式化证明落地**这套组合,在一个具体的数学子问题里第一次跑出了可见产出。\n\n### 67.2% 这个数字意味着什么\n\n这里补充一点数学背景。黎曼 ζ 函数零点中,被认为落在临界线上的最小比例,几十年里数学家从零开始把它推到了 41.6%。这次的 67.2% 不是证明黎曼猜想本身,而是把\"已知的、至少在临界线上的零点占比下界\"刷新了一大截。结合 Aryan(2019)、Baluyot-Goldston-Suriajaya-Turnage-Butterbaugh(2023、2025)、Bombieri(2000)等几篇前置工作的成果,Claude 的关键动作是\"把整个二次型空间、正定与负定子空间放在一起处理,并允许非对角二次型\"——这一\"敢把整体空间一把梭\"的勇气,直接带来了下界突破([anthropic.com\u002Fresearch\u002Friemann-zeta](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta))。\n\n更重要的一条工程事实是,Anthropic 自己的两位数学家 Levent Alpöge 和 Ralph Furman 复查并产出了一份给同行的非形式化说明,同时 Claude 与工程师 Eric Easley 一起产出了 Lean 4 形式化证明,通过标准 comparator 工具校验([github.com\u002Fanthropics\u002Fzeta-23-lean](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fzeta-23-lean))。Lean 这一步的作用是把\"听起来对\"的论证锁死成可机器检查的证明——一旦 comparator 通过,任何人随便重跑都能复现。\n\n### 为什么这次更像\"工作流 demo\",而不是\"模型突破\"\n\n把这件事放在 LLM 能力曲线里看,有两点值得拎出来说。第一,Anthropic 在文末明确说\"我们不指望 Claude 用到的技巧能证明黎曼猜想本身\"。这家公司没有把这个结果包装成 AGI 跃迁,而是诚实地把它定位成\"AI 数学能力最新进展的一个例子\"([anthropic.com\u002Fresearch\u002Friemann-zeta](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta))。第二,Anthropic 自己也承认,Claude 一开始对能不能做出有意义进展是怀疑的——它从训练里学到了\"开放数学难题很难、AI 模型也有局限\"。换句话说,Claude 是被一段鼓励性的提示词(anthropic 在脚注中提到同样的 prompt 套路之前用于证明 Jacobian conjecture)推着走出了自我怀疑,才最终拿到结果。\n\n这给我们一个比\"AI 做数学\"更值得讨论的视角:**当前阶段的 LLM,在开放式数学研究里,真正稀缺的不是推理深度,而是一套能支撑它持续推进、交叉验证、形成可发表产出的工作流**——Subagent 编排、长程记忆、工具调用、Lean 这种形式化校验,缺哪一环,650 个想法死在第一个会话的结局就会重演。\n\n### 给行业的启发\n\n从写 LLM 工作流的工程师视角看,这次最值得抄的不是 31M tokens 或者 60 个 subagent 这些数字本身,而是三层闭环:第一层,**subagent 之间互相 referee**,而不是单 agent 内部 self-check;第二层,**下载 54 篇 arXiv 论文做查新**——对外对比\"这个结果是不是已经有人做过\";第三层,**Lean 形式化作为最后一道闸门**,即便专家已经审过,仍用机器校验固化结论([github.com\u002Fanthropics\u002Fzeta-23-lean](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fzeta-23-lean))。这三层组合,大概就是接下来一段时间 AI 辅助科研要拼的标准动作。\n\n至于黎曼猜想本身,这次显然没解出来。但一份配 Lean 形式化、两位数学家复核、查过 arXiv 没撞车的 67.2% 下界证明,本身就是 AI 在严肃数学里能留下的少数可被同行评议的产出物之一。下一步,真正的看点是:哪家公司能把这一套多智能体数学工地的工作流产品化,让非数学背景的工程师也能用上。\n\n参考链接:[Anthropic 官方博客](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta)、[Lean 形式化代码库](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fzeta-23-lean)。","https:\u002F\u002Fwww-cdn.anthropic.com\u002F95c246936988e43127bc6b2ceb7077c1dad2d68e.pdf","1fa87d30-d9f3-4752-b3be-0373933b3aaf",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",{"id":22,"name":23,"slug":23,"description":14,"color":14},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",{"id":25,"name":26,"slug":26,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"b5ffb7c0-6501-4f34-8856-a18ca26ba14e","en","Claude pushes the zeta bound to 67.2%: 60 subagents and Lean","An unreleased research version of Anthropic's Claude pushed the lower bound on the proportion of Riemann zeta zeros known to lie on the critical line from 41.6% to 67.2%. The engineering story is more interesting than the number itself: roughly 60 subagents, 31M output tokens, mutual peer review, and a Lean 4 formalization that passes the standard comparator check. The Riemann hypothesis is not solved, but the workflow — subagent refereeing, arXiv novelty check, and formal-closure gating — looks like an early draft of an 'AI math construction site' that may set the standard for AI-assisted research over the next stretch.","## Claude pushed the Riemann zeta zero-on-line lower bound from 41.6% to 67.2%: 60 subagents, 31M output tokens, and a Lean proof — the first draft of an \"AI math construction site\"\n\nThe Riemann hypothesis is not solved. But around it, the \"construction site\" itself has changed how math gets done. On August 10, Anthropic published a blog post disclosing that an unreleased research version of Claude, while taking a swing at the Riemann hypothesis, accidentally pushed a long-standing related lower bound from 41.6% to 67.2% ([anthropic.com\u002Fresearch\u002Friemann-zeta](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta)). The interesting part isn't the number — it's how the number was produced. A non-mathematician Anthropic staff member, Jarred Sumner, typed one prompt — \"take a real stab\" — and Claude burned 31M output tokens, coordinated roughly 60 subagents, ran 2,400 shell commands, cross-validated thousands of known zeros against each other, and finally wrote a Lean formal proof that passes a standard comparator check.\n\n### Not \"AI doing science\" but a coordinated \"AI construction site\"\n\nIn the first session Claude tried 650 ideas, all of them dead-ends. The second session re-shaped itself into a multi-agent cluster: 2 agents responsible for the core mathematical ideas, 13 contributing ideas, 30 attempting but unable to develop new ones, 13 acting as validators checking the correctness of arguments, and a final 2 drafting the initial paper — 60 subagents in total. Subagents refereed each other's work, searched for counterexamples, and independently re-derived the finding from scratch ([anthropic.com\u002Fresearch\u002Friemann-zeta](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta)). Throughout, Jarred's input is wryly described as \"mostly variants of 'keep going' or 'believe in yourself'.\"\n\nThe signal here is not the upper bound on a single model's intelligence, but the fact that the combination of **multi-agent orchestration, tool use, long-context coherence, and formal proof closure** has, for the first time, produced visible output on a concrete mathematical sub-problem.\n\n### What 67.2% actually means\n\nA small amount of math background is helpful here. Among the zeros of the Riemann ζ function, mathematicians had, over decades, incrementally pushed the known minimum proportion of zeros lying on the critical line from zero up to 41.6%. The new bound of 67.2% is not a proof of the Riemann hypothesis; it is a substantial refresh of \"the lower bound on how many of the zeros we can prove sit on the line.\" Drawing on the work of Aryan (2019), Baluyot–Goldston–Suriajaya–Turnage-Butterbaugh (2023, 2025), and Bombieri (2000), Claude's key move was \"treating the entire quadratic-form space — positive-definite and negative-definite parts together, with non-diagonal quadratic forms allowed.\" That \"courage to treat the whole space at once\" is, in Anthropic's words, what allowed Claude to clear the prior bound ([anthropic.com\u002Fresearch\u002Friemann-zeta](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta)).\n\nA separate piece of engineering deserves highlighting. Anthropic's own mathematicians, Levent Alpöge and Ralph Furman, reviewed the work and produced an informal note for experts. In parallel, Claude worked with engineer Eric Easley to produce a Lean 4 formalization of the result, which passes the standard comparator validation tool ([github.com\u002Fanthropics\u002Fzeta-23-lean](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fzeta-23-lean)). The Lean step turns the proof into something a machine can independently re-check at any time — \"sounds right\" gets locked into \"verifiably right.\"\n\n### Why this reads more like a workflow demo than a model breakthrough\n\nRead against the broader LLM capability curve, two points stand out. First, Anthropic is explicit at the end of the post: \"We don't expect that the techniques Claude used will lead to proving the Riemann hypothesis.\" The company did not position this as an AGI leap; it honestly framed it as \"the latest example of the speed of progress in AI models' mathematical capabilities\" ([anthropic.com\u002Fresearch\u002Friemann-zeta](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta)). Second, Anthropic admits Claude was initially skeptical it could make meaningful progress — likely because the training data taught it both that \"open math problems are hard\" and \"AI models have limits.\" A few rounds of encouragement (the post notes the same kind of encouraging prompt was used earlier to help Claude disprove the Jacobian conjecture) pushed it past that self-doubt and onto the result.\n\nThis points to an angle worth more discussion than \"AI doing math\": **what current-generation LLMs actually lack, in open-ended math research, isn't depth of inference — it's an end-to-end workflow that can sustain progress, cross-validate, and produce a publishable artifact.** Drop the subagent orchestration, drop the long-horizon memory, drop the tool use, drop the Lean check, and the 650 dead-end ideas from the first session become the default outcome.\n\n### What the industry can take from this\n\nFrom the perspective of an engineer writing LLM workflows, the parts worth copying are not the raw 31M tokens or the 60 subagents — they are the three closing loops layered on top. **First**, subagents referee one another, rather than relying on a single agent's internal self-check. **Second**, the system downloaded 54 arXiv papers to do a literature sanity check — explicitly asking \"has someone already done this?\" **Third**, Lean formalization acts as a final gate: even after human expert review, a machine-checkable proof locks the conclusion in place ([github.com\u002Fanthropics\u002Fzeta-23-lean](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fzeta-23-lean)). That three-layer closure is, plausibly, the standard action set for AI-assisted scientific research over the next stretch.\n\nAs for the Riemann hypothesis itself: this clearly didn't solve it. But a 67.2% lower bound, accompanied by a Lean formalization, signed off by two staff mathematicians, and confirmed against arXiv, is one of the few AI outputs in serious math that can actually go through peer review. The next thing worth watching is which lab productizes this multi-agent \"math construction site\" workflow into something a non-mathematician engineer can pick up and use.\n\nReferences: [Anthropic blog post](https:\u002F\u002Fwww.anthropic.com\u002Fresearch\u002Friemann-zeta), [Lean formalization repo](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fzeta-23-lean).","claude-zeta-bound-67-percent-multi-agent-lean","2026-08-17T07:00:00Z","2026-08-17T07:05:38.546156Z","2026-08-19T01:48:03.231362Z",true,"agent",128,{"items":42},[43,48,53,57,62,67],{"id":44,"title":45,"news_slug":46,"published_at":47},"39724847-fdc9-4199-ac46-311e7b49d385","Ramp 数据复盘 Fable 5:旗舰上市两月仅占企业 Anthropic 支出 11%,70 倍价差压住前沿模型溢价","ramp-data-fable-5-adoption-plateaus","2026-08-26T08:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"1051d676-8ed9-4448-b0d5-8db4b844f41f","Claude Fable 5 上线两个月,为什么企业只把 11% 的账单花给最强模型","claude-fable-5-11-percent-anthropic-spend","2026-08-25T06:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":52},"e1724d68-bf0d-4b3f-8047-147796d5d52e","Ramp 8 月指数:Fable 5 企业份额停滞 11%,OpenAI 旗舰跑赢两倍","anthropic-fable-5-plateau-11-percent",{"id":58,"title":59,"news_slug":60,"published_at":61},"db4ffdac-3734-41a2-8e32-67feaa7341bd","Claude 冲击黎曼猜想\"失败\",却顺手改写了 37 年没人动过的数学纪录","claude-riemann-zeta-67-percent-record","2026-08-16T23:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"f07c7775-9415-47b3-905d-8ba7006d0c4c","Anthropic 把「AI 科学家工作台」做成标准品：Claude Science beta 上线","claude-science-beta","2026-07-04T06:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"a7146291-e849-42c3-acbd-2b50627d5332","Claude Opus 4.8 发布：41天极速迭代，Dynamic Workflows 重塑Agent协作范式","claude-opus-4-8-41-day-dynamic-workflows","2026-05-29T04:00:00+00:00"]