Polylog
The Polylog AI Intelligence Brief

Morning Edition · Sunday, August 2, 2026Published at 1:44 AM EDT · New York

OpenAI Says Unreleased Astra Model Cracked Ten Long-Open Math and Complexity Problems

The results are released with Lean 4 machine-checkable certificates and a 249-page manuscript, with successful proof runs costing roughly $2,000 in tokens.

OpenAI Says Unreleased Astra Model Cracked Ten Long-Open Math and Complexity Problems

OpenAI published ten results it says resolve or advance problems that have seen no progress for at least a decade, spanning high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics. The company attributes the work to an internal version of Astra, described as its next major model.

The claimed advances include the first explicit construction of a non-sofic group, a disproof of Connes's rigidity conjecture, a proof of quantum parallel repetition for general two-player entangled games, a proof of Ehrhart's volume conjecture, and the first improved general sphere-packing exponent since 1978. What separates this from prior lab claims of mathematical ability is verification. OpenAI says the model formalized its arguments in Lean 4, producing machine-checkable certificates on GitHub alongside a 249-page manuscript and per-result reasoning walkthroughs. The company states the successful runs cost roughly $2,000 in tokens at its Sol API rates.

A careful assessment separates two claims. The Lean certificates make the correctness of the formalized proofs independently checkable, which is far stronger evidence than a natural-language argument a reviewer must trust. What remains asserted rather than verified is the process: how many failed attempts, how much human scaffolding shaped the problem setups, and whether the novelty is the model's or comes from expert prompts. Mathematicians will now read the manuscript and audit the Lean files, and their verdict, not the announcement, settles whether these are genuine open-problem resolutions.

Veracity: Plausible
71/100
If true, who benefits

OpenAI's fundraising narrative and valuation, which gain from a frontier-capability claim outside experts can machine-check, pressuring rivals whose reasoning claims rest on unverifiable benchmark numbers.

The nuance

The Lean 4 files make the formalized proofs' correctness independently checkable, but how much human scaffolding shaped each problem setup and whether these are full open-problem resolutions awaits expert audit, and OpenAI overstated a similar Erdős-problems claim in October 2025.

An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.

What this means

If the Lean proofs survive expert audit, the capability that matters is verifiable autonomous reasoning: a frontier lab can attach machine-checkable certificates to research output, which removes the trust gap that has limited AI in mathematics. OpenAI gains a capability claim that outside experts can check rather than one built on selected examples, which pressures rivals to produce formally verified results instead of benchmark scores. The exposed parties are labs whose reasoning claims rest on unverifiable evaluation numbers, and the immediate test is whether working mathematicians reproduce and accept the proofs.

What to watch

  • Independent verification of the Lean 4 certificates by mathematicians outside OpenAI, which would confirm the proofs are correct regardless of how they were produced.
  • Disclosure of the compute and human scaffolding behind each run, which decides whether this is autonomous discovery or heavily guided search.

Observations to monitor, not financial advice.

2 sources

Synthesized from: OpenAI · Polylog editors

Part of a tracked trend

AI Moves Into Autonomous Scientific Discovery and Clinical Care

Over the next 3-9 months, AI systems move beyond text tasks into running real scientific experiments and managing clinical care, backed by peer-reviewed and benchmarked evidence of chemist- and physician-level performance.