Morning Edition · Wednesday, September 9, 2026Published at 2:20 AM EDT · New York
The company says roughly 10,000 parallel agents worked 88 hours to produce a finite-time singularity proof formalized in Lean, while the Clay Mathematics Institute still lists the problem as open.

OpenAI said on September 8 that an unreleased internal system, one it describes as more capable than the GPT-6 Astra model it previewed five days earlier, produced a solution to the Navier-Stokes Millennium Prize Problem. The company published a writeup together with a formal proof in Lean, the proof assistant that mechanically checks each logical step. The claimed result is a finite-time singularity: a smooth fluid flow that breaks down after a finite interval. The Clay Mathematics Institute problem admits either a regularity proof or a demonstration that solutions can fail, so a verified blowup construction would settle it.
The production details are what practitioners should note. Reporting on the announcement puts the run at about 88 hours with as many as 10,000 agents in parallel and a 165-page writeup, a scale of computation closer to a sustained research campaign than a single chat session. Telegram channel AI Post carried the same figures to Russian and English-language AI audiences within hours.
Two things are disputed. First, status: the Clay Mathematics Institute continues to list Navier-Stokes among its unsolved problems, and independent mathematicians have not yet published a confirmation that the released Lean artifact compiles and proves the stated theorem. Second, credit. Tristan Buckmaster of New York University and Levent Alpöge of Anthropic released Lean-verified blowup proofs for related fluid systems, and Buckmaster has alleged that private Codex-session work may have been visible to OpenAI researchers. OpenAI says it finished its own Lean verification on September 6 without accessing that work, and Sébastien Bubeck of OpenAI called the allegations false and inflammatory.
The unusual feature of this dispute is that one half of it is decidable. Originality claims are a matter of testimony, but a Lean file either type-checks under the standard axioms or it does not, and any mathematician with the artifact can run that check.
Part of a tracked trend
AI Moves Into Autonomous Scientific Discovery and Clinical Care
Over the next 3-9 months, AI systems move beyond text tasks into running real scientific experiments and managing clinical care, backed by peer-reviewed and benchmarked evidence of chemist- and physician-level performance.
Start a discussion in Townsquare.
More from this edition
OpenAI, which converts a contested mathematical milestone into evidence that massive parallel inference buys frontier research output, supporting the capital case for continued accelerator spending.
The announcement and the credit dispute are both well documented by Nature, Quanta and Axios, but the article omits the load-bearing detail that the blowup is for the forced Navier-Stokes equations, which several mathematicians treat as a weaker target than the unforced problem, and it omits that OpenAI has said it cannot entirely rule out an indirect connection to Tristan Buckmaster and Levent Alpöge's prior work and began the effort on September 1 after hearing of their result.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Formal verification shifts the constraint on AI-produced mathematics away from years of human refereeing and toward whether a lab releases the machine-checkable artifact. If outside mathematicians compile OpenAI's Lean proof successfully, the claim stands on its own regardless of the credit dispute, and every lab gains an incentive to route hard research claims through proof assistants. If the artifact stays partly withheld, the result remains a vendor assertion, and OpenAI absorbs the reputational cost. Either way the run itself is a demand signal: 10,000 concurrent agents for 88 hours is sustained accelerator consumption for a single question, which supports the case that frontier inference spending grows faster than chat usage alone would justify.
What to watch
Observations to monitor, not financial advice.
Synthesized from: OpenAI · Polylog editors
Comments
0No comments yet.