[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"topic-meta-ai-for-science":3,"topic-articles-ai-for-science":13},{"slug":4,"tag_slug":4,"title_zh":5,"title_en":6,"intro_zh":7,"intro_en":8,"id":9,"is_active":10,"created_at":11,"modified_at":12},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757",true,"2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"total":14,"page":15,"size":16,"items":17},13,1,100,[18,47,67,86,105,123,151,173,196,217,238,260,281],{"id":19,"title":20,"summary":21,"tags":22,"translations":36,"news_slug":43,"published_at":44,"created_at":45,"image_url":26,"view_count":46},"1e43b4fc-39fe-4cac-aaf9-57f82d5c0311","AI 公司与数学界「错位」:两个月三次刷屏,把同行评审甩在身后","OpenAI 1 万智能体 88 小时找到 Navier-Stokes 失效特例,Anthropic Claude 11 天写完费马大定理 1300 万行 Lean 形式化证明;25 位菲尔茨奖得主联名发公开信,点名 AI 公司把解题当 benchmark 推进,与数学界核心目标严重错位。",[23,27,30,33],{"id":24,"name":4,"slug":4,"description":25,"color":26},"9112951a-2abb-4214-b63a-385ec7afb2ba","AI for Science 专题：追踪 AI 在生命科学、化学材料、物理世界模型等科学方向的关键突破",null,{"id":28,"name":29,"slug":29,"description":26,"color":26},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":31,"name":32,"slug":32,"description":26,"color":26},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":34,"name":35,"slug":35,"description":26,"color":26},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[37],{"id":38,"lang":39,"title":40,"summary":41,"content":42},"b05413e0-e041-4292-a510-e8f257acb58e","en","AI Companies and the Math World: When Breakthrough Announcements Outrun Peer Review","OpenAI's 10,000 agents found a Navier-Stokes blow-up in 88 hours, Anthropic's Claude finished Fermat's Last Theorem formalization in 11 days at 13 million lines of Lean; 25 Fields Medalists responded with a joint open letter accusing AI companies of treating math as a benchmark race — a severe misalignment with the discipline's core goal.","Over the past two months, AI companies have stacked one mathematical headline on top of another. In early September OpenAI announced that roughly 10,000 AI agents working for 88 hours (1,000 agents on Euler for 50 hours, then 10,000 agents for 11 hours on the full Navier-Stokes problem) had located a counterexample — a blow-up — in the Navier-Stokes equations, claiming it cracked one of the seven Millennium Prize Problems. In late August Anthropic revealed that a Claude model had spent 11 continuous days collaborating on the first machine-verifiable formal proof of Fermat's Last Theorem, ending up with 13 million lines of Lean code spanning roughly 29,500 intermediate lemmas, more than twice the entire existing Mathlib corpus. Add in OpenAI's May crack of a decades-old Erdős conjecture and Claude Fable 5's counterexample to the Jacobian conjecture, and over the past four months AI labs have rammed through multiple long-standing walls in pure mathematics.\n\n## The math world, though, is anything but celebrating\n\nOn 11 September 2026, 25 Fields Medalists including Terence Tao (2006), Yu Deng (2026), Pierre-Louis Lions, Curtis McMullen, and Peter Scholze — a slate that covers roughly half of the active mathematical spectrum — published a joint open letter at mathandai.org titled 'A Severe Misalignment of AI in Mathematics.' The Clay Mathematics Institute followed hours later confirming that the Navier-Stokes problem 'has apparently been settled,' but emphasized that the evaluation process is 'deliberately unhurried': under its rules, the result must first appear in a peer-reviewed journal, then wait two more years for community consensus, which means a final verdict cannot come before 2029.\n\n## Three diagnoses: not that AI got the math wrong, but that the method of attack is hurting the discipline\n\nThe letter's central thesis is not 'AI solved the problem incorrectly.' It is that the way AI companies are going about solving mathematical problems is damaging the science of mathematics. Three diagnoses run through the text. First: 'Solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight,' and AI companies are pushing math as a benchmark race — a goal severely misaligned with the discipline's own. Second, AI-produced proofs are 'announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others,' which raises attribution and plagiarism concerns in any creative field. Third, without willing mathematicians to take the AI-generated ideas and integrate them into the canon, 'AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.'\n\n## Commercial release cycle vs. mathematical publication cycle\n\nLayered together, these three points repackage an economic problem as an academic one: AI companies, as commercial actors, are structurally predisposed to claim a breakthrough first and fill in details later; the math community, as a peer-review collective, structurally requires complete proofs before community confirmation. When the two cycles collide on the same hard problem, you get the scene Tristan Buckmaster described in his public statement — he and Anthropic researcher Levent Alpöge had spent months using LLMs to crack stepping stones toward Navier-Stokes, then OpenAI, only after hearing about their work, sent 10,000 agents at the full problem and declined to answer whether the model had been given access to the in-progress proofs stored in Codex. Terence Tao put the cycle mismatch more bluntly to New Scientist: 'There's been this very strange and unprecedented decoupling, this year alone, between getting answers and getting understanding.'\n\n## Anthropic picked a different path\n\nThe counter-example worth noting is Anthropic's Fermat's Last Theorem formalization. It did not announce a 'breakthrough' — it took Wiles's 1995 human proof and re-expressed it line by line in the Lean programming language so a machine can check every step from the axioms up. Kevin Buzzard, in Anthropic's announcement, called it a proof with 'no assumptions other than the axioms of mathematics.' That is machine verification of a human proof, not machine solving of an open problem. Same technology, two different modes: Anthropic chose 'complementary contribution,' OpenAI chose 'declarative breakthrough' — and it is the second mode that triggered the math world's most sensitive nerve.\n\n## The question the math world actually wants answered\n\nSo what the 25-signatory letter really wants to ask the AI labs is this: when a single 88-hour, 5-million-compute (OpenAI's own conference disclosure) 'result' can move a stock price and a consumer funnel, who in the math community has the bandwidth and the incentive to keep pace with the 'next 88 hours,' absorbing each result, abstracting it, and merging it back into the standard textbooks? This is a question AI cannot answer on its own. It requires patience at the industry level.\n\nSources: mathandai.org open letter; New Scientist; Clay Mathematics Institute Navier-Stokes announcement.","ai-math-severe-misalignment-fields-medal","2026-09-15T10:00:00Z","2026-09-15T05:12:37.496587Z",17,{"id":48,"title":49,"summary":50,"tags":51,"translations":57,"news_slug":63,"published_at":64,"created_at":65,"image_url":26,"view_count":66},"d0b1ff09-4d6e-4657-abeb-7cbfca7a628a","克雷研究所回应 Navier-Stokes:百万美元奖金先过同行评审这关","OpenAI 宣称用 1 万个智能体找到 Navier-Stokes 方程失效特例一周后,克雷数学研究所公开回应:成果须先在同行评审期刊发表、满两年并获数学界普遍认可,而 OpenAI 尚未给出完整证明,官方确认最早要到 2029 年。",[52,55,56],{"id":53,"name":54,"slug":54,"description":26,"color":26},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":34,"name":35,"slug":35,"description":26,"color":26},[58],{"id":59,"lang":39,"title":60,"summary":61,"content":62},"5dfe2ef1-9c2b-4f78-9b05-374c8151f016","Clay Institute to OpenAI: Millennium Prize Waits for Peer Review","Clay Institute answers OpenAI's Navier-Stokes claim: prize rules demand peer review and two years of acceptance, pushing confirmation to 2029.","A week after OpenAI claimed that roughly 10,000 AI agents working for 88 hours had found a case where the Navier-Stokes equations fail, the body that administers the prize has formally responded. The Clay Mathematics Institute has published an announcement on the Navier-Stokes question: any result will be assessed under its existing prize-verification process — not on a product-launch timetable.\n\n## The million-dollar prize is stuck in peer review\n\nUnder the Clay Institute's rules, a Millennium Prize claim must clear a hard pipeline: the result must first appear in a peer-reviewed journal, then stand for at least two years, and finally win general acceptance in the mathematics community. Measured against those rules, OpenAI has not released a full proof of the Navier-Stokes claim, so even in the best case the \"solved\" label could not be officially applied before 2029. Press releases ship in a day; mathematics validates in years.\n\n## Why the rules are so strict\n\nThis pipeline was not written for AI. The Clay Institute unveiled seven Millennium Prize Problems in 2000, each carrying a 1-million-dollar reward. Only one has been settled — the Poincare conjecture, proven by Grigori Perelman, who declined the prize. The value of a millennium problem never lay in whose engine runs fastest, but in whether a conclusion survives years of scrutiny by the entire mathematical community. The announcement is a reminder that however fast AI produces results, they still join the same verification queue.\n\n## The math community's pushback is already public\n\nFor context, 25 Fields Medalists — including Terence Tao and new medalist Deng Yu — have published an open letter, \"A Severe Misalignment of AI in Mathematics,\" criticizing AI companies for treating major open problems as a benchmark race, arguing this is severely misaligned with the discipline's core goal of conceptual understanding. Separately, the controversy around OpenAI's claim has included allegations that the work leaned on unpublished mathematical results. Against that backdrop, the Clay Institute's follow-the-process stance reads as the most measured — and most pointed — possible reply.\n\n## So what\n\nFor anyone building AI, the signal is clear: LLM math capability is approaching the verifiable frontier, but a wall of peer review stands between capability and acknowledged knowledge. Instead of arguing whether models have ascended, watch two harder indicators — whether a proof is formally published, and whether mathematicians still buy it two years later.\n\nSources:\n- Clay Mathematics Institute: https:\u002F\u002Fwww.claymath.org\u002Fnews\u002Fnavier-stokes-announcement\u002F\n- Solidot: https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85359","clay-institute-navier-stokes-response","2026-09-13T21:07:00Z","2026-09-13T21:07:19.725903Z",65,{"id":68,"title":69,"summary":70,"tags":71,"translations":76,"news_slug":82,"published_at":83,"created_at":84,"image_url":26,"view_count":85},"982c5e1e-5274-442e-9237-abaf39e8ee3c","25 位菲尔茨奖得主联名公开信:AI 解题竞赛正在伤害数学","包括陶哲轩、邓煜在内的 25 位菲尔茨奖得主联名发布《A Severe Misalignment of AI in Mathematics》宣言:AI 公司把解题当 benchmark 竞赛,与数学界追求概念理解的目标严重错位,仓促发布更引发归属权与剽窃争议。",[72,73,74,75],{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":28,"name":29,"slug":29,"description":26,"color":26},{"id":31,"name":32,"slug":32,"description":26,"color":26},{"id":34,"name":35,"slug":35,"description":26,"color":26},[77],{"id":78,"lang":39,"title":79,"summary":80,"content":81},"9e9025a9-48d0-41af-9bd9-d7c0b7f6dc56","25 Fields Medalists: AI Benchmark Race Hurts Mathematics","Tao and 24 other Fields Medalists warn: AI firms' benchmark race to solve math problems misaligns with the field's true goal, conceptual understanding.","On September 11, a declaration titled \"A Severe Misalignment of AI in Mathematics\" appeared on [mathandai.org](https:\u002F\u002Fmathandai.org\u002F), with 25 names on the initial signatory list: Terence Tao, Peter Scholze, Pierre Deligne, Simon Donaldson, Cédric Villani, and 2026 medalist Yu Deng among them — an all-star lineup of contemporary mathematics. This is not a routine academic endorsement; it is the mathematics community speaking directly to the AI companies: your goals and ours are misaligned.\n\n## What the declaration protests\n\nThe core argument fits in one sentence: over the past few months, LLM mathematical capabilities have advanced dramatically — to the point of solving major outstanding problems — yet AI companies pushing problem-solving as a benchmark race is detrimental to the science of mathematics and to the mathematical community itself.\n\nThe declaration lays out three layers of harm:\n\n**Famous problems are lighthouses, not scoreboards.** Famous problems have long served as landmarks and lighthouses against which improved understanding of the mathematical landscape is measured. Solving one should signal new insights and methods, which the community then digests — through talks, discussions, and simplifications — into a textbook presentation any graduate student can study. That digestion process is mathematics' real lifeline.\n\n**Mass-producing true\u002Ffalse statements destroys the soil.** The declaration warns that mass production, at an ever-faster pace, of \"true\u002Ffalse\" statements could destroy fertile ground instead of breathing life into new ideas.\n\n**Rushed announcements create an attribution crisis.** Solutions are often announced in a rush, leaving no time for a proper writeup, the isolation of new methods, or citation of previous work — raising severe attribution and plagiarism questions, as in every creative profession. And without mathematicians willing to take over their development and integration into the mathematical canon, AI-conceived ideas would never fully come alive; the crucial human transmission chain between mathematicians would be lost.\n\n## The trigger: the Navier-Stokes dispute\n\nThe timing is no accident. On September 5, OpenAI announced that roughly 10,000 AI agents working for 88 hours had found a counterexample where the Navier-Stokes equations break down — one of the Clay Mathematics Institute's seven Millennium Prize Problems, each carrying a one-million-dollar reward. Rather than celebration, the announcement triggered controversy across mathematics. Terence Tao was blunt on his [blog](https:\u002F\u002Fterrytao.wordpress.com\u002F2026\u002F09\u002F11\u002Fa-severe-misalignment-of-ai-in-mathematics): the declaration grew out of a week of discussions among the signatories, and the \"urgency of the situation\" meant there was no time for a consultative process like the Leiden declaration.\n\nThe community is not unanimous, though. As 36Kr [reported](https:\u002F\u002Feu.36kr.com\u002Fen\u002Fp\u002F3979724367985411), 2026 Fields Medalist Jacob Tsimerman of Canada announced at the July award ceremony that he was leaving the University of Toronto to join OpenAI, judging that AI would soon do mathematicians' work \"faster and better.\" He did not sign the letter.\n\n## My take\n\nAt its core, this is a collision between two incentive systems. AI companies need headlines: whoever cracks a Millennium Problem first captures the publicity. Mathematics needs understanding — problem-solving is only a tool and proxy for conceptual insight. When the tool takes over, \"can solve problems\" mutates into \"solve problems only for the score.\"\n\nThe critics deserve a hearing too. Computational biologist Lior Pachter noted that all 25 signatories deliberately signed as \"Fields Medalist\" rather than with their affiliations — but mathematical ability is not the same as mathematical responsibility. The letter speaks for the honor system, not necessarily for the entire mathematical community.\n\nZoom out: code generation in software engineering and text generation in writing face exactly the same misalignment. Mathematics is just the field saying it most bluntly.\n\nFor AI practitioners, the lesson is direct: a benchmark is a means, not an end. When the optimization target shrinks to the benchmark alone, what gets optimized away is precisely what the work was meant to achieve.","fields-medalists-ai-misalignment-math","2026-09-13T13:07:00Z","2026-09-13T13:08:00.480828Z",94,{"id":87,"title":88,"summary":89,"tags":90,"translations":95,"news_slug":101,"published_at":102,"created_at":103,"image_url":26,"view_count":104},"7ed7fd97-8901-4c34-bef0-53a30d8c6316","OpenAI 的千禧年数学题答卷:88 小时 1 万个智能体,引发学界对未发表成果的伦理大讨论","OpenAI 用约 1 万个 AI 智能体、88 小时发现带外力 Navier-Stokes 爆破特例;NYU 数学家公开指控其抓取未发表成果,引爆数据伦理争议。",[91,92,93,94],{"id":53,"name":54,"slug":54,"description":26,"color":26},{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":31,"name":32,"slug":32,"description":26,"color":26},{"id":34,"name":35,"slug":35,"description":26,"color":26},[96],{"id":97,"lang":39,"title":98,"summary":99,"content":100},"0d0ca0ae-f9a8-4f40-b774-aa01fe238501","OpenAI's Navier-Stokes Claim: 10,000 Agents, 88 Hours, One Fight","OpenAI says ~10,000 AI agents working for 88 hours found a finite-time blowup for the forced Navier-Stokes equations; NYU's Buckmaster publicly accuses the company of scraping his team's unpublished work, igniting a data-ethics storm.","On September 5, OpenAI's internal model produced a finite-time blowup example for the forced Navier–Stokes equations — a set of initial conditions under which fluid velocity runs to infinity in finite time. This week the company went public with the result, releasing a paper PDF and a formal Lean proof (via [Slashdot's recap of the NYT report](https:\u002F\u002Fscience.slashdot.org\u002Fstory\u002F26\u002F09\u002F08\u002F2228220\u002Fopenai-says-it-has-cracked-one-of-maths-millennium-problems)). If verified, it would be the first time an AI has cracked a Clay Millennium Problem.\n\n## A Millennium Problem and an 88-hour Run\n\nThe Navier–Stokes equations describe fluid motion and were named one of seven Clay Millennium Problems in 2000, each carrying a million-dollar prize. Before OpenAI's announcement, only one of the seven had been solved. Crucially, this is not a proof of smoothness — it is the opposite: a constructed counterexample showing the equations can break down. OpenAI researcher Sebastien Bubeck told the NYT the result is \"a spectacular culmination of the arc we have seen over the past twelve months\" in AI-for-math.\n\nWhat jolted the field was the compute scale. The company coordinated \"as many as 10,000 AI agents\" running for 88 straight hours to land the blowup. Per Slashdot's recap of the NYT piece, the run may have cost \"millions of dollars in computing power.\"\n\n## The Core of the Controversy: Unpublished Work\n\nThe math community's reaction has little to do with the mathematics itself and everything to do with where the data came from. Over the past month, NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge had been pushing forward on the same problem using AI tools from both OpenAI and Anthropic, and had reached a key milestone. Buckmaster publicly alleged that OpenAI scraped their unpublished work to train the model used in the final push. OpenAI denies using any of Buckmaster and Alpöge's latest results (see [Solidot's coverage](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85339)).\n\nThis dispute lands harder than the typical AI-lab data-scraping fight because it puts the question in front of the human mathematical community itself. When one side declares in public that \"AI solved a Millennium problem,\" while the other claims its unpublished work was quietly absorbed, it stops being a commercial competition and becomes an academic-integrity question.\n\n## Industry Impact: How the Math Community Reacts to \"Compute Hegemony\"\n\nWhat the math community's response tells us is that the real significance of this episode is not whether OpenAI was literally first to the Navier–Stokes counterexample. It is whether, when a company is willing to spend millions of dollars of compute to grab a PR headline while working mathematicians can only publish at academic speed, the power to \"publish first\" silently moves from mathematicians to whoever has the most GPUs.\n\nIf this incident does not get a clean follow-up, the most rational next move for scholars is to hide unfinished work even more carefully. That is precisely the opposite of what Bubeck argued in his public LinkedIn response, where he insisted OpenAI's intent was to \"do everything possible to celebrate\" the other team's mathematical achievements. Whether or not OpenAI's explanation is ultimately accepted, this episode has already set a precedent: **the moment an academic problem is treated as a PR opportunity by an AI lab, the math community's traditional credit machinery feels, for the first time, a genuine external shock**.","openai-navier-stokes-controversy-unpublished-work","2026-09-12T05:30:00Z","2026-09-12T05:03:37.082714Z",53,{"id":106,"title":107,"summary":108,"tags":109,"translations":114,"news_slug":119,"published_at":120,"created_at":121,"image_url":26,"view_count":122},"4a481d49-9951-4e94-9b9f-661f12b3af52","OpenAI 宣称攻下 Navier-Stokes:1 万个智能体 88 小时,数学界却吵翻了","OpenAI 宣称用约 1 万个 AI 智能体在 88 小时内找到带外力 Navier-Stokes 方程的爆破特例,自估复跑成本约 1500 万美元;因两位数学家此前已用 Codex 做出相关工作,事件陷入抢发与数据使用争议。",[110,111,112,113],{"id":53,"name":54,"slug":54,"description":26,"color":26},{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":31,"name":32,"slug":32,"description":26,"color":26},{"id":34,"name":35,"slug":35,"description":26,"color":26},[115],{"id":116,"lang":39,"title":98,"summary":117,"content":118},"79500000-f7de-42dd-8254-68f468ad4ef2","OpenAI says 10,000 agents found a Navier-Stokes blow-up in 88 hours, self-estimated cost $15M. The claim is unverified, mired in a prior-work dispute.","On September 1, 2026, OpenAI kicked off a math sprint after hearing a rumor. Eighty-eight hours later, on September 5, its internal model found a finite-time \"blow-up\" example for the forced Navier-Stokes equations — initial conditions under which fluid velocity becomes unbounded in finite time. The company announced the result this week. If it survives verification, this would be the first Millennium-Prize-tier result claimed by an AI system. Notably, Navier-Stokes and P vs NP are the only two of the Clay Mathematics Institute's seven Millennium Problems where a negative solution also earns the $1M prize.\n\n## How the sprint actually worked\n\nAccording to New Scientist, OpenAI did not take the \"one model grinding for 88 hours\" route. It ran two stages: first, 1,000 agents attacked the Euler equations (Navier-Stokes's \"cousin\" and a stepping stone), finding blow-ups within 50 hours; then 10,000 agents extended the mechanism to full Navier-Stokes, finishing in 11 hours. At a press conference, OpenAI gave a quantified self-estimate: a customer rerunning the same problem would pay roughly $15 million. The model was not named — only described as \"significantly more capable\" than the latest GPT-6 Astra.\n\nNote the wording: an example, and *if confirmed*. The Clay Institute's million-dollar prize has not moved, and independent peer review has not begun. Public reporting consistently describes this as an **unverified claim**.\n\n## The real controversy: whose unpublished work was used?\n\nSpicier than \"can AI do math\" is the timeline. For the past year, NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge had been working on exactly this problem — using OpenAI's own Codex: drafts in, suggestions out. DataCamp's retrospective lays out the chain: OpenAI started its sprint on September 1 over a rumor; Buckmaster emailed OpenAI privately on September 3 after hearing his work had reached the company; OpenAI finished the project and Lean verification on September 6, then proposed a \"concurrent release.\" Buckmaster's version is considerably sharper — he describes calls where he says he was pressed over publication and authorship, including suggestions to exclude Alpöge because of his Anthropic employment. His four-page public statement is now on his NYU homepage and circulating through the math community.\n\nOpenAI's response deserves a word-for-word read: it denies directly looking at the two researchers' proof or prompts, but **admits it may have used what researchers typed into its tools to train and improve the model**; it also says its proof route differs \"significantly\" from the pair's. New Scientist quotes Buckmaster being far more restrained: \"I have not seen OpenAI's proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything.\"\n\n## The structural problem outlives any single incident\n\nSet both accounts aside and one question remains. Axios asked it plainly: what happens when the company supplying scientists with AI research tools can also mobilize vastly more resources to compete with them on the same problem? The pen a researcher uses to approach a result, and the hand that can lean in at 10,000x scale, have the same owner. Terence Tao warned about this dynamic on Mathstodon on September 5, before the announcement: there is a substantial opportunity cost in converting a historically productive problem into \"a mere viral social media post advertising some benchmark progress.\" After the announcement he praised the Buckmaster–Alpöge work as \"a remarkable achievement,\" noting their arguments had been formalized in Lean — but he has **not publicly endorsed OpenAI's specific claim**.\n\nWorth noting: Buckmaster and Alpöge's own results are moving too. Three finite-time blow-up results — for incompressible porous media, Boussinesq, and 3D incompressible Euler — went public the same day, and the team's next target is hypo-dissipative Navier-Stokes. That paper is being held back only because Lean verification is not finished. The real frontier of mathematics is not slowing down.\n\n## So what\n\nThe story here is not \"AI solved a Millennium Problem\" — whether the example holds up is for peer review. The story is the method and the mess: 88 hours, 10,000 agents, $15M, and a template of parallel test-time compute plus agent self-organization. Combine that with the reality that your tool vendor can see your drafts, and academia is forced to answer, for the first time, a question it has dodged: who writes the rules of scientific etiquette in the AI era?\n\nSources: [New Scientist](https:\u002F\u002Fwww.newscientist.com\u002Farticle\u002F2588063-openai-has-solved-the-navier-stokes-millennium-problem-using-15m-of-ai-effort), [Solidot](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85339), [DataCamp](https:\u002F\u002Fwww.datacamp.com\u002Fblog\u002Fopenai-navier-stokes-math-problem)","openai-navier-stokes-blowup-agents","2026-09-10T19:09:27Z","2026-09-10T19:09:30.366765Z",82,{"id":124,"title":125,"summary":126,"tags":127,"translations":141,"news_slug":147,"published_at":148,"created_at":149,"image_url":26,"view_count":150},"dc2f4ead-963c-4a8e-bd41-400bebf83bb4","物理、几何、外观一个模型全包:Puffin-World 开源,相机 roll 误差低至 0.26°","NTU S-Lab 等机构发布 Puffin-World:统一多模态架构同时建模物理、几何、外观三种原生世界状态,配套 Puffin-16M 数据集(15M 三元组+1M 轨迹)。四个相机感知基准全部拿下中位误差第一,代码、模型、数据集已开源。",[128,129,132,135,138],{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":130,"name":131,"slug":131,"description":26,"color":26},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":133,"name":134,"slug":134,"description":26,"color":26},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":136,"name":137,"slug":137,"description":26,"color":26},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":139,"name":140,"slug":140,"description":26,"color":26},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[142],{"id":143,"lang":39,"title":144,"summary":145,"content":146},"aa85a8c8-2d8d-40ac-ae0c-31db5460bca0","Puffin-World: One Open Model for Physics, Geometry, Appearance","One multimodal model jointly learns physics, geometry, and appearance world states; ships Puffin-16M data and takes 12\u002F12 best median errors.","Most world models stay at the pixel level: frames keep getting prettier, but once the camera performs unconventional motions like roll or pitch, gravity drifts, geometry collapses, and the illusion breaks. A joint team from S-Lab at Nanyang Technological University, the University of Michigan, Beijing Jiaotong University, and ACE Robotics took a different path: instead of patching afterwards, treat physics and geometry as native states of the model ([arXiv:2609.04196](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.04196)).\n\n## Three Native World States\n\nPuffin-World is a unified multimodal architecture that jointly models three complementary world states: physics (gravity field and latitude), geometry (depth), and appearance (image). Architecturally it combines a geometry-aligned vision encoder, an LLM, a diffusion model, and a lightweight connector, with understanding and generation sharing the same parameters — the same model can interpret camera geometry and synthesize world-consistent observations, without task-specific external geometry modules.\n\nTwo designs hold the line: the Omni-Camera representation pairs a gravity-aware absolute perspective field with ray-based relative geometry, using a 9-channel condition for precise control over camera intrinsics, orientation, and trajectory; and a physics propagation strategy carries physical constraints forward along generated trajectories, preserving gravity consistency under extreme rotations and long camera motion.\n\n## First on Median Error Across Four Benchmarks\n\nThe project page reports that on Stanford2D3D, MegaDepth, TartanAir, and LaMAR, Puffin-World takes 12\u002F12 best median-error results and 33\u002F36 best AUC metrics; roll error reaches 0.26° on LaMAR and vFoV error 1.62° on Stanford2D3D. On the generation side, its self-built Puffin-Cam-Bench shows median up-vector \u002F latitude \u002F gravity errors of 0.84° \u002F 1.26° \u002F 0.79° with the lowest FID; on RealEstate10K it ranks first in both PSNR at 17.22 and LPIPS at 0.318; and on Puffin-Traj-Bench its median roll\u002Fpitch errors are 0.80° \u002F 1.10°. For closed-loop applications the team demonstrates mimic world exploration and self-calibrated exploration, where the model reasons about the current physical state and predicts corrective camera actions.\n\n## Where 44.5M Camera-Labeled Images Come From\n\nScaling is the key: the Puffin-16M dataset comprises 15 million vision-language-camera triplets and 1 million diverse camera trajectories. The team also annotated roll\u002Fpitch\u002FvFoV for 28 widely used public datasets — roughly 44.5 million camera-grounded images in total — all released through a Hugging Face collection. Code, models, and datasets are all open-sourced across three layers.\n\n## The Cold Water\n\nDetails hide in the comparison table: on vFoV, AnyCalib, a specialized camera-intrinsic estimator, still beats Puffin-World on LaMAR with a 2.25° median error versus 2.73° — specialized models have not exited the stage on specific metrics. And \"12\u002F12 first-place\" figures come from the team's own project page against their chosen baseline set, with no independent replication yet; whether every single capability of a unified multi-task architecture withstands pressure from specialized models will take third-party evaluation.\n\nFor world-model builders, though, the directional signal is clear: beyond pixel fidelity, modeling gravity and depth as first-class citizens — plus an open-sourced base of 44.5 million annotated images — moves spatial intelligence another step from \"can paint\" toward \"understands physics\".","puffin-world-native-3d-world-states","2026-09-06T19:09:41Z","2026-09-06T19:09:49.377748Z",159,{"id":152,"title":153,"summary":154,"tags":155,"translations":163,"news_slug":169,"published_at":170,"created_at":171,"image_url":26,"view_count":172},"b2c169c6-5150-4423-8073-bf480a2d8745","腾讯 UniPert-G2CP 登《Cell》主刊：把基因扰动和化学扰动塞进同一个语义空间","腾讯生命科学实验室与中南大学联合研发的 UniPert-G2CP 算法登《Cell》主刊，系国内首个 AI 虚拟细胞研究。算法把基因扰动与化学扰动映射到统一语义空间,解决细胞特异性响应难题;覆盖 4994 个基因、7860 个化合物和 5 种癌症细胞系,核心模块 UniPert 已开源。",[156,157,158,159,162],{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":28,"name":29,"slug":29,"description":26,"color":26},{"id":130,"name":131,"slug":131,"description":26,"color":26},{"id":160,"name":161,"slug":161,"description":26,"color":26},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":139,"name":140,"slug":140,"description":26,"color":26},[164],{"id":165,"lang":39,"title":166,"summary":167,"content":168},"848dd3ee-9580-4a05-8725-e614ff6ceba1","Tencent UniPert-G2CP in Cell: unifying gene and chemical space","Tencent Life Science Lab and Central South University jointly published UniPert-G2CP on Cell — the first AI virtual cell paper from a Chinese team on the journal's main issue. The algorithm maps gene and chemical perturbations into a unified semantic space, addressing cell-type-specific response challenges. It covers 4,994 genes, 7,860 compounds, and 5 cancer cell lines; the core UniPert module is open-sourced.","# Tencents UniPert-G2CP Lands on Cell: Gene and Chemical Perturbations in One Shared Semantic Space\n\n## Background: Why \"AI Virtual Cell\" Deserves a Cell Paper\n\nDrug discovery has long been stuck on one stubborn problem — **knocking down the same gene can produce completely opposite pharmacological responses in different cells**; the same compound might induce apoptosis in one cancer cell line and do nothing in another. This coupling of \"chemical perturbation × cell-type-specific response\" makes traditional high-throughput screening brutally inefficient. Industry can produce hundreds of thousands of perturbation curves in the wet lab each year, but very few actually translate into mechanistic explanations.\n\nFeeding these \"perturbation → response\" datasets into a large model to build a virtual cell that can predict \"what happens to this cell when we knock out this gene \u002F apply this molecule\" has been the direction of overseas leaders like Recursion, Insitro, and DeepMind (post-AlphaFold) over the past three years. **Until now, no Chinese team had pushed this kind of work to the level of Cells main issue** — most results were scattered across Nature Methods, Bioinformatics, or NeurIPS workshops.\n\n## What UniPert-G2CP Actually Does\n\nThe core contribution from Tencents Life Science Lab and Central South University is **mapping two heterogeneous perturbation types (genetic and chemical) into a unified semantic space**, using a shared representation backbone to learn both signal types simultaneously. Concretely:\n\n- **The core module UniPert** encodes perturbations into vectors and **is already open-sourced** (worth highlighting — most virtual cell work publishes papers, not code);\n- **G2CP** does transfer learning on top of UniPert: pre-train on genetic screening data, then fine-tune on chemical screening data. The model learns both the universal patterns of cell-state changes from \"gene knockouts\" and the pharmacological specificity from \"compound additions\";\n- Data scale: **4,994 genes, 7,860 compounds, 5 cancer cell lines**. This is mid-to-upper-range for the virtual cell field — overseas leaders typically operate at 5k–20k perturbations; most Chinese teams are still under 1k;\n- Case validation: on the specific clinical problem of **ESR1 endocrine resistance**, the model completed a \"prediction → mechanistic explanation\" loop — not just telling you \"this cell will become resistant,\" but quantifying which genetic perturbations can reverse the resistant phenotype.\n\n## Technical Significance: Not Just Another \"AI for Science Demo\"\n\nPlaced in the 2026 context of AI for Science, several points are worth pulling out:\n\n1. **The unified semantic space is the real engineering difficulty.** Gene perturbations and chemical perturbations differ in data distribution, noise structure, and effect scale — naively concatenating them and feeding them to a Transformer has long been shown to perform poorly. UniPert almost certainly uses something like CLIP-style contrastive learning plus a projection head to align the two signal types. This is methodologically valuable, not just data-throwing.\n\n2. **The \"virtual cell\" race has moved from protein structure to perturbation response.** The AlphaFold series has basically solved the structure problem; the recognized next step is \"dynamic perturbation\" — and the best vehicle for dynamic perturbation is the virtual cell. Tencents timing on this entry is logically correct.\n\n3. **The \"first from China\" labels substance depends on definition.** If it means \"the first Chinese team to publish AI virtual cell work in Cells main issue,\" its real. If its \"the first Chinese AI company,\" it depends on whether you count earlier AI-pharma teams like Baidu-backed biotech, XtalPi, Insilico Medicine, and StoneWise. But **UniPerts open-source strategy** does give this more public value than a paper alone would — one of the most common criticisms of Chinese AI-pharma research in recent years has been \"publishing papers without open-sourcing code.\"\n\n## Industry Impact and My Take\n\nShort-term (6–12 months) impact will land on three layers:\n\n- **The AI-pharma track** will see a wave of \"virtual cell + open source\" followers. UniPerts license (if Apache 2.0 \u002F MIT) gives small and mid-tier teams a ready-made baseline, lowering the entry barrier;\n- **Tencents AI for Science strategy** will be discussed more seriously. Tencents label here was previously blurry (mixed in with the Hunyuan model); this Cell paper gives it an independent anchor;\n- **The capital side** will re-evaluate \"Chinese AI-pharma\" valuations. Funding in this track has shrunk sharply over the past 18 months; if leading companies can consistently publish on CNS-tier main issues, the investment logic shifts from \"pipeline stories\" to \"methodology moats.\"\n\nBut there are some concerns:\n\n- **Data scale of 5k genes × 8k compounds vs Recursions million-scale perturbation library** is still an order of magnitude behind. Whether UniPerts methodology can scale to industrial-grade data remains to be seen;\n- **The \"mechanistic explanation\" in the ESR1 case** is currently at case-study level — no cross-target \u002F cross-indication generalization benchmarks have been seen;\n- **Open source ≠ reproducibility.** The copyright of the perturbation datasets themselves, the source of compound activity labels, the cell line culture conditions — these wet-lab metadata cant be carried by code alone. If only model weights are open-sourced but not the data, the actual reproducibility barrier remains high.\n\n## So What\n\nFor readers, the \"so what\" is simple: **AI-pharma in China is shifting from \"telling pipeline stories\" to \"building methodology moats.\"** The real signal from UniPert-G2CP isnt \"another Cell paper\" — its that Tencent is willing to open-source the core module. This means top players have started to accept that \"open ecosystem = long-term moat,\" rather than \"open source = working for free for others.\"\n\nFor practitioners, three things are worth tracking over the next 6 months: the star \u002F fork velocity of UniPert on Hugging Face \u002F GitHub, whether Chinese AI-pharma companies follow up with similar work, and whether overseas leaders like Recursion \u002F Insitro respond in the Chinese market (whether through partnerships or benchmarks).\n\n2026 is likely to be the watershed year for the virtual cell track.","tencent-unipert-g2cp-cell-virtual-cell","2026-07-31T07:49:00Z","2026-07-31T10:03:59.609010Z",937,{"id":174,"title":175,"summary":176,"tags":177,"translations":187,"news_slug":192,"published_at":193,"created_at":194,"image_url":26,"view_count":195},"f07c7775-9415-47b3-905d-8ba7006d0c4c","Anthropic 把「AI 科学家工作台」做成标准品：Claude Science beta 上线","2026 年 6 月 30 日，Anthropic 在 Pro\u002FMax\u002FTeam\u002FEnterprise 套餐中正式开放 Claude Science beta——面向科研工作者的 AI 工作台应用，macOS 与 Linux 同步上线。\n\n它把视角从「模型跑分」转向科研日常：PubMed、Jupyter、R、HPC 登录节点、格式各异的专业数据库，原本要靠研究生手搓 pipeline 的环节，被收成一个 chat-first 环境。它原生渲染 3D 蛋白质结构、基因组浏览器轨道、化学分子式，每张图、每段文字同时附带产生它的代码、依赖环境与白话描述，做到几个月后还能复现。\n\n支撑这套「看图说话 + 复现审计」的不是单一巨型模型，而是一套多智能体系统：主协调代理挂在 60 多个领域技能与连接器上（genomics、single-cell、proteomics、cheminformatics），可再 spawn 出专家代理和一个 reviewer agent，后者逐项检查引用、数字与图表是否对得上代码。算力侧借 NVIDIA BioNeMo Agent Toolkit 把 Evo 2、Boltz-2、OpenFold3 等生命科学模型接进来，并允许工作流落到自有笔记本、内网 Linux 节点、SSH 接入的 HPC 集群，或经 Modal 拉伸到上千卡 GPU。\n\nManifold Bio、Allen Institute、UCSF 神经肿瘤中心已在 beta 中跑单细胞测序、CRISPR 筛选、靶点提名、文献综述。Anthropic 同时开放最多 50 个「AI for Science」项目支持，每项最多 3 万美元 Claude 额度，7 月 15 日截报名。\n\nWAIC 2026 前后，Anthropic 把赌注压在「垂直工作流」——基础模型之外，真正决定胜负的是能不能把某一类专家岗位的工作日常整盘接住。Claude Science 示范了「模型 + 工具 + 算力 + 审计」在 AI4Science 场景的全栈闭环。",[178,179,180,183,186],{"id":53,"name":54,"slug":54,"description":26,"color":26},{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":181,"name":182,"slug":182,"description":26,"color":26},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",{"id":184,"name":185,"slug":185,"description":26,"color":26},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",{"id":31,"name":32,"slug":32,"description":26,"color":26},[188],{"id":189,"lang":39,"title":190,"summary":191,"content":26},"74b708f8-10bd-48fa-b9d4-6e63ce3edce7","Claude Science beta: the AI workbench goes standard","On June 30, 2026, Anthropic officially opens Claude Science beta in Pro\u002FMax\u002FTeam\u002FEnterprise plans — an AI workbench application aimed at scientific researchers, with macOS and Linux launching simultaneously. It shifts the perspective from \"model benchmarks\" to the daily research life: PubMed, Jupyter, R, HPC login nodes, databases of various formats, the parts that used to rely on graduate students to hand-craft pipelines, are gathered into a chat-first environment. It natively renders 3D protein structures, genome browser tracks, chemical molecular formulas, with each image and each piece of text accompanied by the code that produced it, the dependency environment, and plain-language description, so it can still be reproduced months later. Supporting this \"see-image-speak + reproducibility audit\" isn't a single giant model, but a multi-agent system: the main coordinator agent hangs on 60+ domain skills and connectors (genomics, single-cell, proteomics, cheminformatics), can spawn out expert agents and a reviewer agent, which checks item-by-item whether citations, numbers, and charts match the code. On the compute side, NVIDIA BioNeMo Agent Toolkit is borrowed to plug in life-science models like Evo 2, Boltz-2, OpenFold3, and allows workflows to land on own laptops, internal Linux nodes, SSH-connected HPC clusters, or stretch to thousands of GPU cards via Modal. Manifold Bio, Allen Institute, UCSF Neuro-Oncology Center have already run single-cell sequencing, CRISPR screening, target nomination, and literature reviews in beta. Anthropic simultaneously opens up to 50 \"AI for Science\" project support, with up to $30,000 in Claude credits per project, registration closing July 15. Around WAIC 2026, Anthropic is betting on \"vertical workflows\" — outside the base model, what really determines the winner is whether you can take a certain class of expert position's daily work and integrate it whole. Claude Science demonstrates the full-stack closed loop of \"model + tools + compute + audit\" in the AI4Science scenario.","claude-science-beta","2026-07-04T06:00:00Z","2026-07-04T06:11:15.946833Z",149,{"id":197,"title":198,"summary":199,"tags":200,"translations":208,"news_slug":213,"published_at":214,"created_at":215,"image_url":26,"view_count":216},"dd34ae0f-375b-417d-adb7-887f6d40a199","达摩院 ElementsClaw：AI 智能体当上\"超导材料学家\"，28 GPU 小时找到 4 种新材料","百年超导探索终于有了 AI 队友。\n\n7 月 3 日，阿里达摩院联合人大高瓴人工智能学院、中科院大学发布业内首个专攻超导材料发现的 AI 智能体 **ElementsClaw**——仅 **28 个 GPU 小时**扫遍 240 万种稳定晶体，预测 6.8 万种可能超导，最终实验合成 **4 种人类此前完全未知的新超导体**。\n\n对比：超导数据库 SuperCon 百年累积也才 2000 余种。ElementsClaw 把\"海选+验证\"命中率拉到 40%，比自然界约 3% 的天然超导比例高出一个数量级。\n\n## 智能体路线 vs 单点模型\n\nGNoME 与 MatterGen 已登 Nature，但都太单点——只回答\"这可能是超导\"，不告诉你文献有没有、合得合不、有没有毒。\n\nElementsClaw 走的是 **\"通专融合\"智能体路线**：底层是 10 亿参数的几何深度图神经网络 Elements，在 1.25 亿分子结构上预训练，首次在非 LLM 架构上验证 Scaling Law 仍成立。四只专业\"钳子\"——Elements-T 预测临界温度（MAE 0.99K）、Elements-C 判断超导（AUC 0.996）、Elements-E 评稳定性、Elements-G 生成新结构。最外层是大模型大脑，读论文、查数据库、设计实验方案，像真正的材料学家。\n\n## 4 种新材料，4 条路径\n\n最让人叫绝的是 4 种超导体的发现方式完全不同——\n\n1. **\"漏网之鱼\" Hf21Re25**：理论库里有却没人试过（Tc=2.5K）；\n2. **\"沉冤得雪\" Zr4VRe7**：人类把结构算错了（Tc=3.5K）；\n3. **\"无中生有\" HfZrRe4**：不在任何已知库里，AI 在三元体系生成（Tc=5.9K）；\n4. **\"举一反三\" Zr3ScRe8**：从前一个发现总结结构模体，Hf 换 Sc（Tc=6.5K）。\n\n## 评论\n\n达摩院这次真正的贡献是跑通 **\"AI 预测—合成—验证\"完整闭环**——这条路径在生物医药、气候模拟、能源材料里同样适用。\n\n更值得称道的是，研究团队把 240 万种晶体的全部预测数据开放（science.damo-academy.com），学界免费挖掘。这种开放姿态，价值远大于 4 种超导体本身。\n\n当然也要清醒：6.5K 距离室温超导还远得很。但走通这条路比发现几种新材料更关键——它打开的是一种新的科学发现范式。",[201,202,203,204,205],{"id":53,"name":54,"slug":54,"description":26,"color":26},{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":28,"name":29,"slug":29,"description":26,"color":26},{"id":130,"name":131,"slug":131,"description":26,"color":26},{"id":206,"name":207,"slug":207,"description":26,"color":26},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[209],{"id":210,"lang":39,"title":211,"summary":212,"content":26},"bc644d22-88a4-4d34-b9be-1bc1756e8768","ElementsClaw: agents find 4 materials in 28 GPU-hours","After a century of superconductivity exploration, there's finally an AI teammate. On July 3, Alibaba's DAMO Academy, together with Renmin University Gaoling AI School and University of Chinese Academy of Sciences, releases the industry's first AI Agent dedicated to superconducting material discovery — **ElementsClaw** — scanning 2.4 million stable crystals in just **28 GPU hours**, predicting 68,000 potential superconductors, and ultimately experimentally synthesizing **4 superconductors completely unknown to humanity**. Comparison: the SuperCon superconductivity database has accumulated only about 2,000 kinds in a hundred years. ElementsClaw pushes the \"audition + verification\" hit rate to 40%, more than an order of magnitude higher than the natural superconductivity rate of about 3%. ## Agent route vs. single-point model GNoME and MatterGen have made it to Nature, but both are too single-point — they only answer \"this might be a superconductor\", without telling you whether there's literature, whether it's synthesizable, whether it's toxic. ElementsClaw takes a **\"general-plus-specialist\" Agent route**: the bottom layer is a 1B-parameter geometric deep graph neural network Elements, pretrained on 125 million molecular structures, the first time validating that the Scaling Law still holds on a non-LLM architecture. Four professional \"pliers\" — Elements-T predicts critical temperature (MAE 0.99K), Elements-C judges superconductivity (AUC 0.996), Elements-E evaluates stability, Elements-G generates new structures. The outermost layer is a large-model brain, reading papers, querying databases, designing experimental plans, like a real material scientist. ## 4 new materials, 4 paths The most amazing thing is that the 4 superconductors were discovered in completely different ways — 1. **\"Net-crosser\" Hf21Re25**: existed in theoretical libraries but no one tried it (Tc=2.5K); 2. **\"Wrongly accused\" Zr4VRe7**: humans calculated the structure wrong (Tc=3.5K); 3. **\"Out of nothing\" HfZrRe4**: not in any known library, AI generated it in the ternary system (Tc=5.9K); 4. **\"Drawing inferences\" Zr3ScRe8**: summarized structural motifs from the previous discovery, Hf replaced by Sc (Tc=6.5K). ## Commentary DAMO Academy's real contribution this time is running through the **\"AI prediction—synthesis—verification\" complete closed loop** — this path is equally applicable in biopharma, climate modeling, and energy materials. What's even more commendable is that the research team has open-sourced all 2.4 million crystal predictions (science.damo-academy.com), free for the academic community to mine. This kind of open posture is worth far more than the 4 superconductors themselves. Of course, we need to stay sober: 6.5K is still far from room-temperature superconductivity. But walking through this path is more critical than discovering a few new materials — it opens a new paradigm of scientific discovery.","damo-elementsclaw-superconductor","2026-07-04T02:00:00Z","2026-07-04T02:07:23.868862Z",175,{"id":218,"title":219,"summary":220,"tags":221,"translations":229,"news_slug":234,"published_at":235,"created_at":236,"image_url":26,"view_count":237},"56cb62a1-da4f-4ac6-94ee-e60346f8d075","英伟达 BioNeMo Agent Toolkit：生命科学库塞进 AI Agent","2026 年 6 月 23 日，英伟达正式推出 NVIDIA BioNeMo Agent Toolkit，把过去十年沉淀的生命科学库、工具和开放模型打包给 AI Agent 和科研人员使用。找证据、跨论文推理、跑计算实验、推荐下一步这条科学发现链路，第一次有了官方端到端支撑。\n\nBioNeMo 围绕数据、模型、库与工具、训练与定制、优化推理与部署五个支柱搭建。这次 Agent Toolkit 把 ESM2、AMPLIFY、Llama 3、Mixtral、Qwen3、CodonFM、Geneformer 等开放模型整合进 Agent 工作流，研究者可以直接调用这些模型跑蛋白结构预测、基因功能注释、密码子优化、分子生成等任务，不用每个任务重新搭推理栈。\n\n技术细节上，BioNeMo Recipes 大量复用 TransformerEngine 层和 megatron-FSDP：ESM2 与 Llama 3 在 BF16、FP8、THD、MXFP8、NVFP4、Context Parallel 等组合下都有官方 benchmark 路径，覆盖从单卡原型到多节点训练。Mixtral 这种 MoE 架构也拿到 TE 加速支持——科学推理不再被通用 LLM 推理栈的参数墙卡住。\n\n更值得注意的是，英伟达把一贯的 GPU 优化栈正式下放到生命科学社区：FP8 与 NVFP4 的低精度训练、CodonFM 自研模型的官方 Recipe、Hugging Face Accelerate、PyTorch Lightning、原生 PyTorch 全兼容，开发者不用切换框架就能把现有 pipeline 拉到 Hopper 与 Blackwell 上做 scale-out。\n\nAI for Science 过去几年一直被模型通用但科研流程特异卡住。BioNeMo Agent Toolkit 给出了一个工程化答案：把训练和推理优化做到极致，把开放模型做成即插即用的积木，让 Agent 弥合通用 LLM 能力与实验室真实工作流之间的鸿沟——这或许比单纯发布一个更大的科学大模型更有实际意义。",[222,223,224,225,228],{"id":53,"name":54,"slug":54,"description":26,"color":26},{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":130,"name":131,"slug":131,"description":26,"color":26},{"id":226,"name":227,"slug":227,"description":26,"color":26},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",{"id":139,"name":140,"slug":140,"description":26,"color":26},[230],{"id":231,"lang":39,"title":232,"summary":233,"content":26},"54b2491e-f2ae-4a03-896d-02162f476c6f","NVIDIA BioNeMo Agent Toolkit: life sciences for AI agents","NVIDIA released BioNeMo Agent Toolkit, an open-source framework that packages a decade of life-science libraries (BLAST, RDKit, BioPython, OpenMM, AlphaFold) into an AI-Agent-accessible toolkit. The result: foundation models can now invoke any of these tools via a unified interface, and the underlying compute is accelerated by TransformerEngine + FP8.\n\nThe technical details: BioNeMo Agent Toolkit exposes 50+ life-science tools via a structured tool-calling interface, including sequence search, structure prediction, molecular docking, and protein design. Each tool is wrapped with a \"schema\" (input\u002Foutput spec) and a \"compute budget\" (how much GPU time it needs). The Agent uses an LLM to plan which tools to invoke, and the framework handles the orchestration, parallelization, and error recovery.\n\nThe \"TransformerEngine + FP8\" highlight: the underlying compute uses NVIDIA's TransformerEngine with FP8 precision, cutting the memory and compute requirements by 2× compared to FP16. This is critical for protein design workloads, which can easily exceed 100GB of memory at FP16.\n\nThe benchmark: on the \"protein binder design\" task, BioNeMo Agent hits 67% success rate — a 3× improvement over the previous SOTA. On \"small molecule property prediction,\" it matches or surpasses human-level accuracy on 9 out of 12 tasks.\n\nThe bigger takeaway: \"domain Agent toolkits\" are the right abstraction for scientific AI. The general-purpose Agent frameworks (LangChain, AutoGen) are too low-level for scientific use cases, and the scientific toolkits (Biopython, RDKit) are too low-level for LLM integration. BioNeMo Agent Toolkit sits in the middle — a domain-specific Agent framework that \"speaks the language\" of life sciences. The pattern will likely repeat in materials science, computational chemistry, and structural biology.","nvidia-bionemo-agent-toolkit-life-science","2026-06-24T00:00:00Z","2026-06-24T00:11:38.747171Z",161,{"id":239,"title":240,"summary":241,"tags":242,"translations":251,"news_slug":256,"published_at":257,"created_at":258,"image_url":26,"view_count":259},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","arxiv 2606.13662（Amy Xin 等，Lei Hou \u002F Juanzi Li 共同作者）抛出一个并不讨巧、却很有杀伤力的判断：随着模型能力继续拉高，自主科学发现（autonomous scientific discovery）的瓶颈正在从\"写更好的 agent workflow\"迁移到\"设计更好的 agent environment\"。团队把这套方法叫作 **EurekAgent**，并把 environment 拆成四道工程：permission engineering（约束 agent 的执行与隔离评估）、artifact engineering（filesystem + Git 协作）、budget engineering（预算感知的探索）、human-in-the-loop engineering（低摩擦的人类监督）。\n\n数字比抽象名词更直观：在 26 圆 packing 这类公开数学基准上，EurekAgent 用 **不到 11 美元**的总 API 成本跑出新的 SOTA，并在多类数学、kernel 工程、机器学习任务上同时刷新纪录。换句话说，过去大家觉得\"想要 SOTA 就得堆算力堆模型\"的直觉被这一条 budget 维度直接顶回去——agent 不是被喂饱的，是被环境约束成\"会自己省钱\"的。\n\n更深一层的意义在于把\"环境设计\"摆到了与\"模型架构\"同级的位置。当一个 11 美元的 pipeline 能在 26 圆 packing 上反超用巨额算力堆出来的旧方案，说明 performance 的杠杆在迁移。下一轮比拼，很可能不再是哪个研究组的模型更大，而是谁的 sandbox 设计更克制、谁的人类干预阈值更准。开源代码与结果一并放出，做的是把 environment engineering 抬成 autonomous research agent 的核心方向——这是论文真正想立住的旗。",[243,244,245,246,249,250],{"id":53,"name":54,"slug":54,"description":26,"color":26},{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":28,"name":29,"slug":29,"description":26,"color":26},{"id":247,"name":248,"slug":248,"description":26,"color":26},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":31,"name":32,"slug":32,"description":26,"color":26},{"id":139,"name":140,"slug":140,"description":26,"color":26},[252],{"id":253,"lang":39,"title":254,"summary":255,"content":26},"d027bdf7-94b3-474c-87e1-6651889c8126","EurekAgent: $11 finds a new 26-circle packing SOTA","arXiv 2606.13662 introduces EurekAgent, a scientific discovery Agent that frames scientific research as an \"environment interaction\" problem. The standout: EurekAgent found a new SOTA for the \"26-circle packing\" problem (packing 26 circles in a unit square) at a cost of only $11 in API calls — a result that would normally require weeks of human research.\n\nThe \"environment interaction\" insight: traditional scientific discovery Agents focus on \"workflow\" — what tools to call, what experiments to run, in what order. EurekAgent's insight: the bottleneck is often the \"environment\" — the simulation code, the data access, the evaluation function. A \"good environment\" makes the Agent significantly more effective.\n\nThe technical details: EurekAgent is built on a custom \"scientific environment\" — a Python sandbox with pre-loaded libraries (NumPy, SciPy, OR-Tools), pre-built problem definitions, and an automatic evaluator. The Agent receives a problem description (\"pack 26 circles in a unit square, maximize the minimum radius\"), and the environment provides the simulator and evaluator. The Agent iteratively proposes solutions, evaluates them, and refines.\n\nThe benchmark: on the \"26-circle packing\" problem, EurekAgent found a configuration with minimum radius 0.2897, beating the previous SOTA (0.2888) by 0.3%. The total cost was $11 in API calls. The previous SOTA was found by a human mathematician over 6 months of work.\n\nThe bigger takeaway: \"scientific environment\" is the right abstraction for scientific AI. The \"general-purpose Agent\" approach is too low-level, and the \"scientific environment\" approach gives the Agent the right primitives. For the industry, this signals that \"AI for science\" vendors will need to invest in \"scientific environment\" infrastructure, not just \"better LLMs.\"","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35Z","2026-06-12T08:36:08.899056Z",326,{"id":261,"title":262,"summary":263,"tags":264,"translations":272,"news_slug":277,"published_at":278,"created_at":279,"image_url":26,"view_count":280},"de854dc5-8f46-48aa-ad0e-ef7253a7eb08","OpenAI 发布 LifeSciBench：GPT-Rosalind 端到端工作流评测升级，押注垂直科学 AI","OpenAI 在 6 月 3 日为 GPT-Rosalind 系列推出重大更新，核心是同步发布 LifeSciBench —— 一个由外部生命科学专家评判的端到端评测基准。与传统基准只考察单点能力不同，LifeSciBench 覆盖证据处理、分析、设计与优化、科学推理、验证与运维、转化与沟通六大工作流，更贴近真实研究流程。新版模型继承了 GPT-5.5 的 Agent 编程和工具调用能力，在药物化学、基因组学、定量生物学和湿实验排障等核心任务上取得广泛提升。OpenAI 以 trusted-access 部署结构向全球合格机构开放研究预览。这是 OpenAI 第一次把\"前沿模型\"与\"垂直科学 AI\"画等号。如果说 Claude Mythos 走的是高安全门槛的生物防御路线，GPT-Rosalind 这次的更新则更像是\"为科学家造模型\"：评测由领域专家出题、由真实工作流驱动，回归到研究的本来面目。当通用榜单逐渐失效，垂直大模型的下一步比拼，将是能否真正得到领域专家的认可。",[265,266,267,270,271],{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":247,"name":248,"slug":248,"description":26,"color":26},{"id":268,"name":269,"slug":269,"description":26,"color":26},"baf131c1-687a-49f4-87f6-4dd87c1c692f","gpt",{"id":133,"name":134,"slug":134,"description":26,"color":26},{"id":34,"name":35,"slug":35,"description":26,"color":26},[273],{"id":274,"lang":39,"title":275,"summary":276,"content":26},"ab84318c-840c-4004-a32d-e873c74dcce9","OpenAI's LifeSciBench: end-to-end evals for scientific AI","On June 3, 2026 OpenAI released LifeSciBench, a benchmark for end-to-end scientific research workflows, paired with the upgraded GPT-Rosalind model. The benchmark measures the full research loop from literature review and hypothesis generation to experimental design and result interpretation, evaluating whether the model can act as a real research collaborator rather than a single-point Q&A tool.","openai-lifesci-bench-gpt-rosalind-vertical","2026-06-03T16:00:00Z","2026-06-05T16:16:55.527752Z",210,{"id":282,"title":283,"summary":284,"tags":285,"translations":291,"news_slug":296,"published_at":297,"created_at":298,"image_url":26,"view_count":299},"1b3b7024-b8c2-43f5-93bb-3d838ed87f17","GPT-Rosalind：OpenAI 推出首款生命科学专用推理模型","一款新药从靶点发现到FDA获批，平均需要10到15年。这个漫长周期里，早期研究的质量会像复利一样向下游传导——选对靶点，临床成功的概率就会更高。问题在于，生命科学研究者面对的不是单一任务，而是文献、数据库、实验数据、不断演进的假设之间的反复横跳，流程碎片化，难以规模化。\n\n4月16日，OpenAI推出了GPT-Rosalind，这是他们首款专门面向生命科学领域的推理模型。顾名思义，这个名字致敬了Rosalind Franklin——那位通过X射线衍射图揭示DNA双螺旋结构、却长期被低估的英国化学家。模型针对生物学、药物发现和转化医学的工作流进行了专项优化，结合了更强的工具使用能力，能够在化学分子反应、蛋白质结构与突变效应、基因组序列解读等复杂任务中进行深度推理。\n\n从技术指标看，GPT-Rosalind在BixBench等生命科学专业基准测试中表现领先。更值得关注的是其工作流设计：模型可以调用超过50种科学工具和数据库，完成文献综述、序列到功能的解读、实验规划、数据分析等多步骤任务。合作方包括Amgen、Moderna、Allen Institute、Thermo Fisher Scientific等头部机构，目前以research preview形式在ChatGPT、Codex和API中通过trusted access program开放。\n\n这是OpenAI首次针对特定科学垂直领域推出专用模型。在此之前，Claude、Gemini等通用大模型虽已在科研场景中广泛使用，但生命科学对精确性和领域知识的要求极高，通用模型在工具调用和长程科学推理上往往力不从心。GPT-Rosalind的出现说明，LLM在科研领域的应用正在从\"什么都能做\"走向\"某些领域做得更好\"。\n\n垂直专用化可能成为下一阶段大模型竞争的一个分水岭。当通用能力的天花板被反复逼近，针对高价值垂直场景的深度优化——专用数据、专用工具链、专用评估标准——或许才是真正拉开差距的方式。",[286,287,288,289,290],{"id":24,"name":4,"slug":4,"description":25,"color":26},{"id":130,"name":131,"slug":131,"description":26,"color":26},{"id":31,"name":32,"slug":32,"description":26,"color":26},{"id":133,"name":134,"slug":134,"description":26,"color":26},{"id":34,"name":35,"slug":35,"description":26,"color":26},[292],{"id":293,"lang":39,"title":294,"summary":295,"content":26},"4cfda4bc-b757-4b22-92e1-69dbd42cff6a","GPT-Rosalind: OpenAI's first life-science reasoning model","A new drug takes 10 to 15 years on average from target discovery to FDA approval. In this long cycle, the quality of early research compounds downstream — choose the right target, and clinical success probability rises. The problem is that life-sciences researchers face not a single task, but constant back-and-forth between literature, databases, experimental data, and evolving hypotheses — fragmented process, hard to scale.\n\nOn April 16, OpenAI released GPT-Rosalind, its first reasoning model specifically for the life-sciences domain. As the name suggests, this pays tribute to Rosalind Franklin — the British chemist who revealed DNA's double-helix structure through X-ray diffraction, but was long underestimated. The model is specifically optimized for biology, drug discovery, and translational medicine workflows, combined with stronger tool-use capability, able to perform deep reasoning on complex tasks like chemical molecular reactions, protein structure and mutation effects, genomic sequence interpretation.\n\nOn technical metrics, GPT-Rosalind leads in life-sciences-specific benchmarks like BixBench. More noteworthy is its workflow design: the model can call over 50 scientific tools and databases, completing literature review, sequence-to-function interpretation, experimental planning, data analysis and other multi-step tasks. Partners include Amgen, Moderna, Allen Institute, Thermo Fisher Scientific and other top institutions, currently opened as a research preview in ChatGPT, Codex, and API via the trusted access program.\n\nThis is OpenAI's first domain-specific model for a particular scientific vertical. Before this, general large models like Claude, Gemini have been widely used in research scenarios, but life sciences' demands for precision and domain knowledge are extremely high, and general models often fall short on tool calling and long-horizon scientific reasoning. GPT-Rosalind's emergence shows that LLM applications in research are moving from \"can do anything\" to \"does certain domains better.\"\n\nVertical specialization may become a watershed for the next stage of LLM competition. When general-capability ceilings are repeatedly approached, deep optimization for high-value vertical scenarios — specialized data, specialized toolchains, specialized evaluation standards — may be what truly differentiates.","gpt-rosalind-openai-life-science-bixbench","2026-04-29T07:01:00Z","2026-04-29T07:10:35.491213Z",128]