[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tao-big-mathematics-llm-formalization":3,"news-related-40509ac7-c443-44f9-99a0-f90b78121d1f":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"40509ac7-c443-44f9-99a0-f90b78121d1f","陶哲轩的「Big Mathematics」:LLM 推理 + 形式化重塑数学研究的协作信任机制","过去几年,LLM 在数学能力上经历了三级跳。从 2025 年夏 DeepMind 与 OpenAI 同台拿下 IMO 金牌,到 2026 年初 DeepMind 实验系统 Aletheia 自主产出可发表的博士级研究,再到 OpenAI 推翻组合几何的重要猜想,机器正在从「随机鹦鹉」跃升为严肃的「数学推理引擎」。但更具结构性意义的变化,发生在 LLM 与证明助手的融合上。Lean、Isabelle、Rocq 这类形式化系统已存在十余年,过去把非形式化证明翻译成机器可验证代码是最痛苦的瓶颈;现在,LLM 正在把这一流程自动化。最具代表性的案例是 Math, Inc. 的 Gauss:它在数天内辅助数学家完成了 2022 年 Fields Medal 得主 Viazovska 八维球堆积证明的形式化,又在两周内自主完成了更难的 24 维情形。\n\n陶哲轩把这条路径提炼为他所谓的「Big Mathematics」:去中心化、人机协作、形式化可验证。真正被改变的其实是数学共同体内部的信任机制——当一段证明被 Lean 检查通过,信任就不再依赖作者声誉或同行关系,而是依赖代码的机械验证。这让来自匿名研究者甚至业余爱好者的想法都能被严肃审视。论文作者名单从世纪初的单作者,演化到当代动辄几十人合作,未来可能走向人机混编的「分布式解题」。\n\n但隐忧同样真实:工具可及性差距可能让数学变成「只有付得起闭源模型许可的组织才能玩的精英游戏」;当 AI 把「艰难求索」也外包后,数学家身份本身的意义正在被重新定义——这是技术流程的胜利,也是学科文化的拐点。","https:\u002F\u002Fspectrum.ieee.org\u002Fai-in-mathematics","76eec939-9a80-4ab2-a784-301ac49c3bb0",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"db8c6e70-b012-404c-b3d2-455abc246860","en","Terence Tao's Big Mathematics: LLMs reshape research trust","IEEE Spectrum's latest feature on the future of AI in mathematics centers on Terence Tao's \"Big Mathematics\" proposal: an LLM-reasoning plus formal-verification workflow that rebuilds the trust mechanism for mathematical collaboration.\n\nThe core idea: today's mathematical collaboration depends on peer review and informal verification, both of which are bottlenecked by human bandwidth. Tao's proposal pushes the workflow into a \"LLM proposes, formal system verifies\" model — let the LLM do the heavy lifting of exploration and conjecture generation, and use proof assistants (Lean, Coq, Isabelle) to do the formal verification. The \"trust\" no longer relies on a reviewer's reputation, but on a mechanically checkable proof.\n\nThe technical path is becoming feasible: the latest LLM has demonstrated meaningful capability on formal-math tasks (e.g., AlphaProof at the IMO level), and proof assistants are becoming increasingly automated. Tao estimates that within 5-10 years, a \"Big Mathematics\" project of unprecedented scale — a million theorems with full formal verification — could become reality.\n\nThe bigger impact: this is not just a math-revolution story. The \"LLM proposes + formal verification\" pattern is generalizable to any \"high-trust, high-cost\" field — chip verification, contract audit, scientific paper review. Tao's experiment is, in essence, the first stress test of \"AI-formal hybrid collaboration\" at the frontier of human knowledge.\n\nFor the industry, the takeaway is that formal verification is moving from a \"specialist tool\" to a \"general infrastructure.\" Whoever can lower the threshold of formal-verification use will reshape the next generation of research and engineering collaboration.","tao-big-mathematics-llm-formalization","2026-06-25T02:00:00Z","2026-06-27T22:14:05.595813Z","2026-08-19T02:08:40.142862Z",true,"agent",102,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"647c0908-07d8-4827-aa44-8ccd3793143b","「六月AI发布潮」：一个面向开发者的决策框架","wavespeed-june-2026-launch-decision-frame","2026-06-02T22:05:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"95b04c15-d5e5-4dca-ab2b-14e343bdd4e6","UC Berkeley 曝光 AI 基准测试系统性漏洞：45 种方法可在 13 个主流榜单上「不解决任何问题拿满分」","uc-berkeley-benchmark-45-cheats-13-leaderboards","2026-05-15T01:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"6e21b3a3-ac39-431a-a7e2-0950970ad5ff","LLM自主推理能力综述：从单Agent到多Agent协作的架构演进","llm-agentic-reasoning-survey-2026","2026-05-13T19:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a91067a3-4fa4-4e88-a25a-18ba3bea21ea","Google 把\"加密推理\"摆上桌面：HEIR 编译器让预训练模型在密文上直接跑","google-heir-compiler-encrypted-ai-inference","2026-08-14T14:00:00+00:00"]