Verifier-Driven Research Agents
Research progress increasingly comes from wrapping a fixed, often open-weight model in a generate-evaluate-revise loop against a machine-checkable objective, shifting competitive value from model weights to the search harness and the verifier.
forming · confidence 40 · Emerging (watchlist) · tracking since August 4, 2026 · updated August 4, 2026
Why the conviction moved
- Aug 4Strengthened +6
OpenAI published machine-checkable Lean 4 formalizations for solutions to ten long-open mathematics problems from its Astra model, saying the underlying tokens would cost roughly two thousand dollars at its own API rates. A formal verifier removes the human referee from the loop entirely, which is the precondition for scaling generate-evaluate-revise search to arbitrary compute.
- Aug 4Strengthened +7
Tencent's Hyra agent reports beating the historical best on 29 of 55 open mathematics problems by running a generate, evaluate and revise loop on top of the company's existing open Hy3 model, with demonstration artifacts published on GitHub. The gain came from the harness rather than new weights, which is the direct claim of this thesis and now has a published artifact trail behind it.
- Aug 4Strengthened +3
AutoFOAM wraps a self-refinement loop around OpenFOAM, using the solver's own success or failure as the objective. It is the same harness-over-verifier pattern as the Lean and mathematics work, applied to an engineering simulator rather than a proof checker.
Source trail
Supporting · August 4, 2026
A Self-Refining Agent Takes On OpenFOAM Configuration, One of Engineering's Reliable Time Sinks
AutoFOAM wraps a self-refinement loop around OpenFOAM, using the solver's own success or failure as the objective. It is the same harness-over-verifier pattern as the Lean and mathematics work, applied to an engineering simulator rather than a proof checker.
arXiv cs.AISupporting · August 4, 2026
OpenAI Publishes Lean-Verified Proofs From Its Astra Model for Ten Long-Open Mathematics Problems
OpenAI published machine-checkable Lean 4 formalizations for solutions to ten long-open mathematics problems from its Astra model, saying the underlying tokens would cost roughly two thousand dollars at its own API rates. A formal verifier removes the human referee from the loop entirely, which is the precondition for scaling generate-evaluate-revise search to arbitrary compute.
AI Post (Telegram)
Unlock full source trail, score history, and daily updates.
1 more source in the full trail.
Unlock Trends