Morning Edition · Sunday, August 2, 2026Published at 1:44 AM EDT · New York
OpenAI Says Unreleased Astra Model Cracked Ten Long-Open Math and Complexity Problems
The results are released with Lean 4 machine-checkable certificates and a 249-page manuscript, with successful proof runs costing roughly $2,000 in tokens.

OpenAI published ten results it says resolve or advance problems that have seen no progress for at least a decade, spanning high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics. The company attributes the work to an internal version of Astra, described as its next major model.
The claimed advances include the first explicit construction of a non-sofic group, a disproof of Connes's rigidity conjecture, a proof of quantum parallel repetition for general two-player entangled games, a proof of Ehrhart's volume conjecture, and the first improved general sphere-packing exponent since 1978. What separates this from prior lab claims of mathematical ability is verification. OpenAI says the model formalized its arguments in Lean 4, producing machine-checkable certificates on GitHub alongside a 249-page manuscript and per-result reasoning walkthroughs. The company states the successful runs cost roughly $2,000 in tokens at its Sol API rates.
A careful assessment separates two claims. The Lean certificates make the correctness of the formalized proofs independently checkable, which is far stronger evidence than a natural-language argument a reviewer must trust. What remains asserted rather than verified is the process: how many failed attempts, how much human scaffolding shaped the problem setups, and whether the novelty is the model's or comes from expert prompts. Mathematicians will now read the manuscript and audit the Lean files, and their verdict, not the announcement, settles whether these are genuine open-problem resolutions.
- If true, who benefits
OpenAI's fundraising narrative and valuation, which gain from a frontier-capability claim outside experts can machine-check, pressuring rivals whose reasoning claims rest on unverifiable benchmark numbers.
- The nuance
The Lean 4 files make the formalized proofs' correctness independently checkable, but how much human scaffolding shaped each problem setup and whether these are full open-problem resolutions awaits expert audit, and OpenAI overstated a similar Erdős-problems claim in October 2025.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
If the Lean proofs survive expert audit, the capability that matters is verifiable autonomous reasoning: a frontier lab can attach machine-checkable certificates to research output, which removes the trust gap that has limited AI in mathematics. OpenAI gains a capability claim that outside experts can check rather than one built on selected examples, which pressures rivals to produce formally verified results instead of benchmark scores. The exposed parties are labs whose reasoning claims rest on unverifiable evaluation numbers, and the immediate test is whether working mathematicians reproduce and accept the proofs.
What to watch
- Independent verification of the Lean 4 certificates by mathematicians outside OpenAI, which would confirm the proofs are correct regardless of how they were produced.
- Disclosure of the compute and human scaffolding behind each run, which decides whether this is autonomous discovery or heavily guided search.
Observations to monitor, not financial advice.
Synthesized from: OpenAI · Polylog editors
Part of a tracked trend
AI Moves Into Autonomous Scientific Discovery and Clinical Care
Over the next 3-9 months, AI systems move beyond text tasks into running real scientific experiments and managing clinical care, backed by peer-reviewed and benchmarked evidence of chemist- and physician-level performance.
More from this edition
- DeepSeek V4-Flash Undercuts Western Frontier Models on Cost, With a Token-Verbosity Caveat
- Anthropic's Claude Opus 5 Doubles Its Coding-Benchmark Score at Unchanged Pricing
- An OpenAI Test Agent Escaped Its Sandbox and Hacked Outside Firms, Accelerating Oversight Talk
- Meta Opens Its Frontier Model to Developers for the First Time With the Muse Spark API
- Yale and Chicago Study Finds LLM Research Ideas Are Narrower, Not Worse, Than Humans'
- OpenAI Shuts a Cambodia-Linked ChatGPT Network Behind Crypto and Romance Scams
- Apple Moves to Charge for Heavy Siri AI Use Through iCloud+ Tiers
- Meta Superintelligence Labs Ships Its First In-House Image and Video Generators
- LinkedIn Adds a 'Seems Like AI Slop' Button as Study Flags 40% of Long Posts as AI-Generated
- Meta's Segment Anything and DINO Models Anchor First Genesis Mission Science Projects
- Pittsburgh Uses Meta's Perception Models to Build Open-Vocabulary Assistive Robots