Morning Edition · Tuesday, September 8, 2026Published at 2:18 AM EDT · New York
Anthropic says a team of Claude agents worked largely autonomously for 11 days and proved 29,500 intermediate theorems, a formalization more than five times the size of Lean's main mathematics library.

Anthropic published what it describes as the first complete, computer-checked proof of Fermat's Last Theorem, written in the Lean proof assistant by Claude agents over 11 days with limited human intervention. The result is large: roughly 13 million lines of Lean code and 29,500 intermediate theorems, more than five times the size of Mathlib, the community-maintained Lean mathematics library.
The important property is not the size but the checking. Anthropic says the proof uses only Lean's three standard axioms and contains no omitted steps, meaning no placeholder arguments and no additional axioms introduced to close a gap. That claim can be verified mechanically by anyone who runs the Lean kernel against the published repository, which puts it in a different category from a self-reported benchmark score. Kevin Buzzard of Imperial College London, who leads the long-running human effort to formalize the same theorem in Lean, called it an extraordinary autoformalization achievement.
Two limits deserve stating plainly. The underlying mathematics was already proved by Andrew Wiles and Richard Taylor in the 1990s, so this is a translation of an existing argument into machine-checkable form rather than a new theorem. And the human blueprint work that broke the proof into formalizable pieces predates the model. What changed is the speed of a task where correctness can be checked by a machine, which is exactly the setting where long-horizon agents are least likely to produce an undetected error.
Anthropic, which converts a machine-checkable artifact into evidence for long-horizon agent reliability, and the vendors selling formal-verification capacity that would absorb the resulting demand.
Part of a tracked trend
AI Moves Into Autonomous Scientific Discovery and Clinical Care
Over the next 3-9 months, AI systems move beyond text tasks into running real scientific experiments and managing clinical care, backed by peer-reviewed and benchmarked evidence of chemist- and physician-level performance.
Start a discussion in Townsquare.
More from this edition
Kevin Buzzard says he compiled the repository successfully yet also says the result adds essentially no new mathematical content, and the Lean kernel checks that the logic follows from its axioms without checking that each of the 29,500 lemma statements means what its name implies, a semantic gap the "fully verified" framing does not carry.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Formal verification is the clearest case where an agent's output can be validated without having to trust the agent, and that is the channel through which this result matters commercially. The same process, generating a candidate proof and then having a checker reject or accept it, applies to hardware design rules, cryptographic protocol proofs, compiler correctness and safety-critical software. Vendors selling verification tools gain a large new source of demand for proof capacity, while the scarce resource shifts from proof engineers to computing power and to the blueprints that break a problem into checkable pieces.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Anthropic Research · Anthropic Research (paper)
Comments
0No comments yet.