The Polylog AI Intelligence Brief

Morning Edition · Tuesday, August 4, 2026Published at 2:03 AM EDT · New York

OpenAI Publishes Lean-Verified Proofs From Its Astra Model for Ten Long-Open Mathematics Problems

Each result ships with a machine-checkable Lean 4 formalization, and OpenAI says the tokens behind all ten would have cost roughly two thousand dollars at its own API rates.

OpenAI Publishes Lean-Verified Proofs From Its Astra Model for Ten Long-Open Mathematics Problems

OpenAI published ten results on August 1 that it says an internal version of its next model family, named Astra, produced in mathematics, operator algebras, quantum complexity and theoretical computer science. The release includes a 249-page technical manuscript, a 62-page account of how the arguments came together, and a formalization of each result in the proof-checking language Lean 4, posted on GitHub. That last element is the one that matters. A Lean proof either type-checks or it does not, which moves these claims out of the category of assertions a lab makes about its own model and into the category of artifacts anyone can run.

The named results are substantial. According to OpenAI's writeup, Astra produced the first explicit construction of a non-sofic group, a question open since the mathematician Mikhail Gromov introduced the notion of soficity in 1999. It also disproved Connes's rigidity conjecture on von Neumann algebras, proved Ehrhart's volume conjecture, and settled three problems from the catalogue of the mathematician Paul Erdős, including the entry on multicoloured Ramsey numbers. OpenAI says each of the ten problems had been open for at least a decade.

The economics attached to the claim are as notable as the mathematics. OpenAI states that the tokens used to generate all ten solutions would have cost about two thousand dollars at its Sol API rates. If that figure holds, the marginal cost of a research-grade proof attempt has fallen to an amount a single graduate student could expense, which changes who can afford to run a search at this scale.

Separate and much weaker claims are circulating alongside this one. Emad Mostaque, the former chief executive of Stability AI, said in remarks relayed on Telegram that a billion-parameter model trained on nothing published after 1911 independently reconstructed general relativity, and that artificial intelligence had recovered more than a century of missing algebra in Einstein's equations. No paper, weights or replication accompany that assertion, and it should be read as a claim by an interested party rather than as a result. The distinction between the two is exactly the distinction Lean certificates are designed to enforce.

Veracity: Corroborated
84/100
If true, who benefits

OpenAI, which converts a capability claim into an artifact competitors cannot answer with self-reported scores, timed shortly before Astra becomes a commercial product, and every vendor selling research-tier inference at premium rates.

The nuance

A Lean 4 file that type-checks proves the formal statement was derived correctly, not that the formalization faithfully encodes the conjecture specialists care about, no peer review has been completed, and the roughly two-thousand-dollar figure counts tokens in the successful runs at OpenAI's own list prices, excluding failed searches, training and the human direction that framed the problems.

An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.

What this means

Formal verification changes the burden of proof for capability claims in a way benchmarks never did, because a Lean 4 file can be checked adversarially by anyone with a laptop and the compiler. That gives OpenAI a defensible marketing asset that Chinese and open-weight competitors cannot match with self-reported scores, and it puts pressure on every lab making reasoning claims to ship machine-checkable artifacts instead of evaluation tables. The exposure runs to research tooling and to the mathematics community itself, where the reviewing bottleneck moves from whether a proof is correct to whether the problem was worth posing. What is verified here is that ten Lean files compile. What is asserted, and not yet independently assessed, is that the underlying problems are as significant as OpenAI says.

What to watch

  • Whether working mathematicians in group theory and operator algebras confirm that the non-sofic group construction and the Connes rigidity disproof are what OpenAI describes. Endorsement from specialists, or a correction from them, decides how much of this survives.
  • Whether other labs start attaching formal certificates to reasoning claims. If Anthropic, Google DeepMind or the Chinese labs follow, machine-checkable output becomes the new minimum standard for a frontier reasoning announcement.
  • When Astra becomes generally available and at what price, since the two-thousand-dollar figure is only meaningful if outside researchers can run the same search themselves.

Observations to monitor, not financial advice.

3 sources

Synthesized from: Polylog editors · OpenAI · The Decoder

Part of a tracked trend

AI Moves Into Autonomous Scientific Discovery and Clinical Care

Over the next 3-9 months, AI systems move beyond text tasks into running real scientific experiments and managing clinical care, backed by peer-reviewed and benchmarked evidence of chemist- and physician-level performance.

Share this article

Comments

0

No comments yet.