Morning Edition · Tuesday, August 4, 2026Published at 2:03 AM EDT · New York
OpenAI Publishes Lean-Verified Proofs From Its Astra Model for Ten Long-Open Mathematics Problems
Each result ships with a machine-checkable Lean 4 formalization, and OpenAI says the tokens behind all ten would have cost roughly two thousand dollars at its own API rates.
OpenAI published ten results on August 1 that it says an internal version of its next model family, named Astra, produced in mathematics, operator algebras, quantum complexity and theoretical computer science. The release includes a 249-page technical manuscript, a 62-page account of how the arguments came together, and a formalization of each result in the proof-checking language Lean 4, posted on GitHub. That last element is the one that matters. A Lean proof either type-checks or it does not, which moves these claims out of the category of assertions a lab makes about its own model and into the category of artifacts anyone can run.
The named results are substantial. According to OpenAI's writeup, Astra produced the first explicit construction of a non-sofic group, a question open since the mathematician Mikhail Gromov introduced the notion of soficity in 1999. It also disproved Connes's rigidity conjecture on von Neumann algebras, proved Ehrhart's volume conjecture, and settled three problems from the catalogue of the mathematician Paul Erdős, including the entry on multicoloured Ramsey numbers. OpenAI says each of the ten problems had been open for at least a decade.
The economics attached to the claim are as notable as the mathematics. OpenAI states that the tokens used to generate all ten solutions would have cost about two thousand dollars at its Sol API rates. If that figure holds, the marginal cost of a research-grade proof attempt has fallen to an amount a single graduate student could expense, which changes who can afford to run a search at this scale.
Separate and much weaker claims are circulating alongside this one. Emad Mostaque, the former chief executive of Stability AI, said in remarks relayed on Telegram that a billion-parameter model trained on nothing published after 1911 independently reconstructed general relativity, and that artificial intelligence had recovered more than a century of missing algebra in Einstein's equations. No paper, weights or replication accompany that assertion, and it should be read as a claim by an interested party rather than as a result. The distinction between the two is exactly the distinction Lean certificates are designed to enforce.
- If true, who benefits
OpenAI, which converts a capability claim into an artifact competitors cannot answer with self-reported scores, timed shortly before Astra becomes a commercial product, and every vendor selling research-tier inference at premium rates.
- The nuance
A Lean 4 file that type-checks proves the formal statement was derived correctly, not that the formalization faithfully encodes the conjecture specialists care about, no peer review has been completed, and the roughly two-thousand-dollar figure counts tokens in the successful runs at OpenAI's own list prices, excluding failed searches, training and the human direction that framed the problems.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Formal verification changes the burden of proof for capability claims in a way benchmarks never did, because a Lean 4 file can be checked adversarially by anyone with a laptop and the compiler. That gives OpenAI a defensible marketing asset that Chinese and open-weight competitors cannot match with self-reported scores, and it puts pressure on every lab making reasoning claims to ship machine-checkable artifacts instead of evaluation tables. The exposure runs to research tooling and to the mathematics community itself, where the reviewing bottleneck moves from whether a proof is correct to whether the problem was worth posing. What is verified here is that ten Lean files compile. What is asserted, and not yet independently assessed, is that the underlying problems are as significant as OpenAI says.
What to watch
- Whether working mathematicians in group theory and operator algebras confirm that the non-sofic group construction and the Connes rigidity disproof are what OpenAI describes. Endorsement from specialists, or a correction from them, decides how much of this survives.
- Whether other labs start attaching formal certificates to reasoning claims. If Anthropic, Google DeepMind or the Chinese labs follow, machine-checkable output becomes the new minimum standard for a frontier reasoning announcement.
- When Astra becomes generally available and at what price, since the two-thousand-dollar figure is only meaningful if outside researchers can run the same search themselves.
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · OpenAI · The Decoder
Part of a tracked trend
AI Moves Into Autonomous Scientific Discovery and Clinical Care
Over the next 3-9 months, AI systems move beyond text tasks into running real scientific experiments and managing clinical care, backed by peer-reviewed and benchmarked evidence of chemist- and physician-level performance.
More from this edition
- Alibaba Ships Qwen3.8-Max, a 2.4-Trillion-Parameter Model It Promises to Open-Weight Next Week
- Tencent's Hyra Agent Reports Beating the Historical Best on 29 of 55 Open Mathematics Problems
- An Open-Source Runtime Puts an 80-Billion-Parameter Qwen Model on a Mac in 4.3 Gigabytes of Memory
- Meta Puts Muse Image and a Muse Video Preview Into Its Consumer Apps With Conversational Editing
- Meta's Open Vision Models Move Into a $41.5 Million Robotic Wheelchair Program and Berkeley Lab Science
- AI Evaluation Turns Into a Product Market as Researchers Show Cheap Open Models Can Grade Proofs
- New Benchmarks Target the Decision Agents Make Before They Answer
- Agent Observability Reaches Cloud Consoles as Enterprises Report Measured Deployment Results
- Researchers Propose a Zero-Trust Registry for Agent Skills After Finding Claims Do Not Match Capabilities
- A Self-Refining Agent Takes On OpenFOAM Configuration, One of Engineering's Reliable Time Sinks
- A Paper Revives Leibniz, Turing and Searle to Argue Consciousness Tests Belong in AI Safety Work
Comments
0No comments yet.