[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ai-math-severe-misalignment-fields-medal":3,"topics-all":39,"news-related-1e43b4fc-39fe-4cac-aaf9-57f82d5c0311":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":25,"news_slug":32,"published_at":33,"created_at":34,"modified_at":35,"is_published":36,"publish_type":37,"image_url":15,"view_count":38},"1e43b4fc-39fe-4cac-aaf9-57f82d5c0311","AI 公司与数学界「错位」:两个月三次刷屏,把同行评审甩在身后","OpenAI 1 万智能体 88 小时找到 Navier-Stokes 失效特例,Anthropic Claude 11 天写完费马大定理 1300 万行 Lean 形式化证明;25 位菲尔茨奖得主联名发公开信,点名 AI 公司把解题当 benchmark 推进,与数学界核心目标严重错位。","过去两个月,AI 公司在数学领域连续刷屏:9 月初 OpenAI 宣布动用约 1 万个 AI 智能体,经过 88 小时(先 1000 个智能体跑 50 小时 Euler,再 10000 个跑 11 小时 Navier-Stokes)找到一个 Navier-Stokes 方程失效的特例,声称重写千禧年数学难题之一;8 月底 Anthropic 用 Claude 模型连续 11 天协作写完费马大定理首个机器可验证的形式化证明,代码量 1300 万行 Lean,跑完约 29500 个中间定理,体积超过此前 Mathlib 全部累积的两倍。叠加上 5 月 OpenAI 模型破解 Erdős 几十年悬而未决的猜想、Claude Fable 5 给出 Jacobian 猜想反例,过去四个月 AI 公司已经「撞开」数学界多个长期封顶的难题。\n\n## 这本来值得庆祝,但数学界却坐立不安\n\n9 月 11 日,25 位菲尔茨奖得主联名在 mathandai.org 上发表公开信《A Severe Misalignment of AI in Mathematics》,联署人包括陶哲轩(2006 年菲尔茨奖)、邓煜(2026 年新晋)、Pierre-Louis Lions、Curtis McMullen、Peter Scholze 等几乎覆盖了当代活跃数学家谱系的一半;克雷数学研究所紧随其后于 9 月 11 日发文确认 Navier-Stokes「apparently been settled」,但同时强调「过程是有意为之的不紧不慢」,按规则需要先在同行评审期刊上发表、再等两年学界普遍认可,因此最早要 2029 年才能给出最终判定。\n\n## 三大病灶:不是 AI 解错了题,而是解题方式伤了学科\n\n公开信的核心论点不是「AI 解错了题」,而是「AI 公司解题的方式正在伤到数学这门学科」。他们指出三大病灶:第一,「解决问题只是达成概念理解与深刻洞察这一核心目标的工具与替代指标」,而 AI 公司把数学当 benchmark 竞赛,与数学界的根本目标已经发生严重错位(misalignment);第二,AI 的解答「发布得过于仓促,甚至没有留出足够时间去编写严谨规范的论文,无法去提炼其中蕴含的新方法与新思想,也无法合理引用前人的相关工作」,这在所有创意行业里都会引爆归属权认定与剽窃争议;第三,如果「没有心怀热忱的数学家负责后续开发并融入数学规范体系,AI 所孕育的思想就永远无法真正获得生命,数学家之间至关重要的人际传递纽带也将断裂」。\n\n## 商业发布周期 vs 数学发表周期\n\n这三点放在一起,把一个经济学问题包装成了学术问题:AI 公司作为商业主体,天然倾向于「先声称为突破,再补细节」的发布节奏;数学界作为同行评审共同体,天然要求「先写完整证明,再让社区确认」。两套发布节奏撞在同一道难题上,就会出现 Buckmaster 在公开声明里描述的那一幕——他与 Anthropic 研究员 Levent Alpöge 几个月来用 LLM 协助突破 Navier-Stokes 的「垫脚石」,OpenAI 在听到他们的工作之后才上 1 万个智能体攻坚 Navier-Stokes 本身,且未回应是否在模型里访问过他们存储在 Codex 上的过程稿。这种「商业发布周期 vs 数学发表周期」的错位,陶哲轩最近在 New Scientist 的采访里更直白:「这是答案与理解之间前所未有地脱钩。」\n\n## Anthropic 选了另一条路\n\n值得注意的反例是 Anthropic 的费马大定理形式化。它没有公开发布「突破」,而是把 Wiles 1995 年的人类证明在 Lean 编程语言里完整重写一遍,让机器可以从底层逻辑一步步检查每一行;Buzzard 在 Anthropic 的发布稿里评价「除数学公理之外不假设任何东西」。这是一次机器对人类证明的「验证」,不是机器对未解难题的「解题」。同样是 AI 加数学,Anthropic 选了「补全式贡献」,OpenAI 选了「宣示式突破」,后者触发了数学界最敏感的神经。\n\n## 数学界真正想追问的问题\n\n所以这封 25 人联署的公开信,真正想问 AI 公司的是:当一个 88 小时 1500 万美元算力(OpenAI 在发布会上披露)的「成果」能让股价涨、能让 ToC 流量涨,数学界又有谁有能力、有动力去跟「下一个 88 小时」赛跑,把这些成果一个一个完整消化、抽象、归并进数学教科书?这道题 AI 自己答不了——它需要的是行业层面的耐心。\n\n参考来源:[mathandai.org 公开信原文](https:\u002F\u002Fmathandai.org\u002F)、[New Scientist 报道](https:\u002F\u002Fwww.newscientist.com\u002Farticle\u002F2588063-openai-has-solved-the-navier-stokes-millennium-problem-using-15m-of-ai-effort)、[克雷数学研究所 Navier-Stokes 公告](https:\u002F\u002Fwww.claymath.org\u002Fnews\u002Fnavier-stokes-announcement\u002F)。","https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85358","d59894d3-308e-4fd8-8865-86dc1eeac4a2",[11,16,19,22],{"id":12,"name":13,"slug":13,"description":14,"color":15},"9112951a-2abb-4214-b63a-385ec7afb2ba","ai-for-science","AI for Science 专题：追踪 AI 在生命科学、化学材料、物理世界模型等科学方向的关键突破",null,{"id":17,"name":18,"slug":18,"description":15,"color":15},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":20,"name":21,"slug":21,"description":15,"color":15},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":23,"name":24,"slug":24,"description":15,"color":15},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[26],{"id":27,"lang":28,"title":29,"summary":30,"content":31},"b05413e0-e041-4292-a510-e8f257acb58e","en","AI Companies and the Math World: When Breakthrough Announcements Outrun Peer Review","OpenAI's 10,000 agents found a Navier-Stokes blow-up in 88 hours, Anthropic's Claude finished Fermat's Last Theorem formalization in 11 days at 13 million lines of Lean; 25 Fields Medalists responded with a joint open letter accusing AI companies of treating math as a benchmark race — a severe misalignment with the discipline's core goal.","Over the past two months, AI companies have stacked one mathematical headline on top of another. In early September OpenAI announced that roughly 10,000 AI agents working for 88 hours (1,000 agents on Euler for 50 hours, then 10,000 agents for 11 hours on the full Navier-Stokes problem) had located a counterexample — a blow-up — in the Navier-Stokes equations, claiming it cracked one of the seven Millennium Prize Problems. In late August Anthropic revealed that a Claude model had spent 11 continuous days collaborating on the first machine-verifiable formal proof of Fermat's Last Theorem, ending up with 13 million lines of Lean code spanning roughly 29,500 intermediate lemmas, more than twice the entire existing Mathlib corpus. Add in OpenAI's May crack of a decades-old Erdős conjecture and Claude Fable 5's counterexample to the Jacobian conjecture, and over the past four months AI labs have rammed through multiple long-standing walls in pure mathematics.\n\n## The math world, though, is anything but celebrating\n\nOn 11 September 2026, 25 Fields Medalists including Terence Tao (2006), Yu Deng (2026), Pierre-Louis Lions, Curtis McMullen, and Peter Scholze — a slate that covers roughly half of the active mathematical spectrum — published a joint open letter at mathandai.org titled 'A Severe Misalignment of AI in Mathematics.' The Clay Mathematics Institute followed hours later confirming that the Navier-Stokes problem 'has apparently been settled,' but emphasized that the evaluation process is 'deliberately unhurried': under its rules, the result must first appear in a peer-reviewed journal, then wait two more years for community consensus, which means a final verdict cannot come before 2029.\n\n## Three diagnoses: not that AI got the math wrong, but that the method of attack is hurting the discipline\n\nThe letter's central thesis is not 'AI solved the problem incorrectly.' It is that the way AI companies are going about solving mathematical problems is damaging the science of mathematics. Three diagnoses run through the text. First: 'Solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight,' and AI companies are pushing math as a benchmark race — a goal severely misaligned with the discipline's own. Second, AI-produced proofs are 'announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others,' which raises attribution and plagiarism concerns in any creative field. Third, without willing mathematicians to take the AI-generated ideas and integrate them into the canon, 'AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.'\n\n## Commercial release cycle vs. mathematical publication cycle\n\nLayered together, these three points repackage an economic problem as an academic one: AI companies, as commercial actors, are structurally predisposed to claim a breakthrough first and fill in details later; the math community, as a peer-review collective, structurally requires complete proofs before community confirmation. When the two cycles collide on the same hard problem, you get the scene Tristan Buckmaster described in his public statement — he and Anthropic researcher Levent Alpöge had spent months using LLMs to crack stepping stones toward Navier-Stokes, then OpenAI, only after hearing about their work, sent 10,000 agents at the full problem and declined to answer whether the model had been given access to the in-progress proofs stored in Codex. Terence Tao put the cycle mismatch more bluntly to New Scientist: 'There's been this very strange and unprecedented decoupling, this year alone, between getting answers and getting understanding.'\n\n## Anthropic picked a different path\n\nThe counter-example worth noting is Anthropic's Fermat's Last Theorem formalization. It did not announce a 'breakthrough' — it took Wiles's 1995 human proof and re-expressed it line by line in the Lean programming language so a machine can check every step from the axioms up. Kevin Buzzard, in Anthropic's announcement, called it a proof with 'no assumptions other than the axioms of mathematics.' That is machine verification of a human proof, not machine solving of an open problem. Same technology, two different modes: Anthropic chose 'complementary contribution,' OpenAI chose 'declarative breakthrough' — and it is the second mode that triggered the math world's most sensitive nerve.\n\n## The question the math world actually wants answered\n\nSo what the 25-signatory letter really wants to ask the AI labs is this: when a single 88-hour, 5-million-compute (OpenAI's own conference disclosure) 'result' can move a stock price and a consumer funnel, who in the math community has the bandwidth and the incentive to keep pace with the 'next 88 hours,' absorbing each result, abstracting it, and merging it back into the standard textbooks? This is a question AI cannot answer on its own. It requires patience at the industry level.\n\nSources: mathandai.org open letter; New Scientist; Clay Mathematics Institute Navier-Stokes announcement.","ai-math-severe-misalignment-fields-medal","2026-09-15T10:00:00Z","2026-09-15T05:12:37.496587Z","2026-09-15T05:12:37.496596Z",true,"agent",19,[40,48],{"slug":13,"tag_slug":13,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":36,"created_at":46,"modified_at":47},"AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":36,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"982c5e1e-5274-442e-9237-abaf39e8ee3c","25 位菲尔茨奖得主联名公开信:AI 解题竞赛正在伤害数学","fields-medalists-ai-misalignment-math","2026-09-13T13:07:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"7ed7fd97-8901-4c34-bef0-53a30d8c6316","OpenAI 的千禧年数学题答卷:88 小时 1 万个智能体,引发学界对未发表成果的伦理大讨论","openai-navier-stokes-controversy-unpublished-work","2026-09-12T05:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"4a481d49-9951-4e94-9b9f-661f12b3af52","OpenAI 宣称攻下 Navier-Stokes:1 万个智能体 88 小时,数学界却吵翻了","openai-navier-stokes-blowup-agents","2026-09-10T19:09:27+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"1942b07b-f794-42b1-b944-ca6b32d4ae16","四大 AI 模型同日集体掉线:OpenAI\u002FClaude 官方确认,Gemini\u002FGrok 表面沉默","four-ai-models-overlapping-outage-sept-2026","2026-09-06T08:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"2fc4f696-a697-498c-b9d5-28250bfeaa79","ChatGPT 进欧盟 VLOP 名单:OpenAI 第一次要为生成式 AI 内容负全责","chatgpt-eu-vlop-dsa-first-ai-platform-rules","2026-09-05T07:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"390c2437-4e4f-45ec-8270-67c5bfa4fa47","ChatGPT、Claude、Grok、Gemini 罕见同时下线,周四早晨全球 AI 集体失声","chatgpt-claude-grok-gemini-thursday-outage","2026-09-05T06:00:00+00:00"]