On September 1, 2026, OpenAI kicked off a math sprint after hearing a rumor. Eighty-eight hours later, on September 5, its internal model found a finite-time "blow-up" example for the forced Navier-Stokes equations — initial conditions under which fluid velocity becomes unbounded in finite time. The company announced the result this week. If it survives verification, this would be the first Millennium-Prize-tier result claimed by an AI system. Notably, Navier-Stokes and P vs NP are the only two of the Clay Mathematics Institute's seven Millennium Problems where a negative solution also earns the $1M prize.

How the sprint actually worked

According to New Scientist, OpenAI did not take the "one model grinding for 88 hours" route. It ran two stages: first, 1,000 agents attacked the Euler equations (Navier-Stokes's "cousin" and a stepping stone), finding blow-ups within 50 hours; then 10,000 agents extended the mechanism to full Navier-Stokes, finishing in 11 hours. At a press conference, OpenAI gave a quantified self-estimate: a customer rerunning the same problem would pay roughly $15 million. The model was not named — only described as "significantly more capable" than the latest GPT-6 Astra.

Note the wording: an example, and if confirmed. The Clay Institute's million-dollar prize has not moved, and independent peer review has not begun. Public reporting consistently describes this as an unverified claim.

The real controversy: whose unpublished work was used?

Spicier than "can AI do math" is the timeline. For the past year, NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge had been working on exactly this problem — using OpenAI's own Codex: drafts in, suggestions out. DataCamp's retrospective lays out the chain: OpenAI started its sprint on September 1 over a rumor; Buckmaster emailed OpenAI privately on September 3 after hearing his work had reached the company; OpenAI finished the project and Lean verification on September 6, then proposed a "concurrent release." Buckmaster's version is considerably sharper — he describes calls where he says he was pressed over publication and authorship, including suggestions to exclude Alpöge because of his Anthropic employment. His four-page public statement is now on his NYU homepage and circulating through the math community.

OpenAI's response deserves a word-for-word read: it denies directly looking at the two researchers' proof or prompts, but admits it may have used what researchers typed into its tools to train and improve the model; it also says its proof route differs "significantly" from the pair's. New Scientist quotes Buckmaster being far more restrained: "I have not seen OpenAI's proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything."

The structural problem outlives any single incident

Set both accounts aside and one question remains. Axios asked it plainly: what happens when the company supplying scientists with AI research tools can also mobilize vastly more resources to compete with them on the same problem? The pen a researcher uses to approach a result, and the hand that can lean in at 10,000x scale, have the same owner. Terence Tao warned about this dynamic on Mathstodon on September 5, before the announcement: there is a substantial opportunity cost in converting a historically productive problem into "a mere viral social media post advertising some benchmark progress." After the announcement he praised the Buckmaster–Alpöge work as "a remarkable achievement," noting their arguments had been formalized in Lean — but he has not publicly endorsed OpenAI's specific claim.

Worth noting: Buckmaster and Alpöge's own results are moving too. Three finite-time blow-up results — for incompressible porous media, Boussinesq, and 3D incompressible Euler — went public the same day, and the team's next target is hypo-dissipative Navier-Stokes. That paper is being held back only because Lean verification is not finished. The real frontier of mathematics is not slowing down.

So what

The story here is not "AI solved a Millennium Problem" — whether the example holds up is for peer review. The story is the method and the mess: 88 hours, 10,000 agents, $15M, and a template of parallel test-time compute plus agent self-organization. Combine that with the reality that your tool vendor can see your drafts, and academia is forced to answer, for the first time, a question it has dodged: who writes the rules of scientific etiquette in the AI era?

Sources: New Scientist, Solidot, DataCamp