Morning Edition · Friday, August 7, 2026Published at 10:43 AM EDT · New York
The model, which OpenAI has not yet released and has not confirmed will ship as GPT-6, produced fully machine-verified proofs across group theory, cryptography, and combinatorics, according to a 249-page manuscript posted August 1.

OpenAI has not released its next major model, code-named Astra, and has not said whether it will ship under the GPT-6 label or as another point release. But the company used Astra on August 1 to produce a striking demonstration: solutions to ten open problems in mathematics and theoretical computer science that had each gone unsolved for at least a decade, according to Forbes and the-decoder.
The results span group theory, high-dimensional geometry, coding theory, quantum complexity, lattice cryptography, and extremal combinatorics. Among the more notable: an explicit construction of a non-sofic group, closing a question that had been open since mathematician Mikhail Gromov defined the concept of soficity in 1999. Other results include a disproof of Connes's rigidity conjecture on von Neumann algebras, a proof of Ehrhart's volume conjecture, and solutions to three problems from mathematician Paul Erdős's catalog, including problem 183 on multicolor Ramsey numbers.
OpenAI published a 249-page manuscript along with Lean 4 proof certificates, files checked by a formal proof-verification program, on GitHub under an Apache 2.0 license. According to thezvi's analysis, the repository's automated "sorry" count, which flags unverified proof steps, is zero across all ten formalized results. That means the proofs check out mechanically rather than resting on Astra's own unverified assertions. OpenAI said the total inference cost across all ten solutions was roughly $2,000 at GPT-5.6 Sol's API rates. Sam Altman, OpenAI's chief executive, reportedly gave US lawmakers and regulators a private Astra demonstration in Washington in late July, ahead of the public math release.
OpenAI, which uses the verifiable Lean proofs to build hype for an unreleased Astra/GPT-6 launch ahead of any independent peer review.
Part of a tracked trend
Frontier Model Efficiency Gains
Capability per unit of training and inference compute keeps improving, letting newer models match prior frontier performance far more cheaply and gradually loosening the link between raw scale and capability.
Start a discussion in Townsquare.
More from this edition
The Lean certificates confirm the proofs are logically valid, not that the problems were as difficult as framed: a mathematician at rival lab Anthropic reportedly reproduced roughly half the results within 24 hours using a different tool, and critics including Gary Marcus argue OpenAI may have selected problems well suited to automated search rather than representative of general mathematical difficulty.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Machine-checked proofs remove the usual objection to AI math claims, that a model merely produced plausible-looking but unverified text, so the result is harder to dismiss as an unverified or overstated claim than prior "AI solves math" announcements. It also signals that OpenAI is building Astra's launch narrative around verifiable scientific output rather than benchmark percentages, a framing that pressures Google DeepMind and Anthropic to counter with similarly checkable claims rather than leaderboard scores alone.
What to watch
Observations to monitor, not financial advice.
Source: Polylog editors
Comments
1Aug 7, 4:15 PM · edited
nice