← Trends

Verifier-Driven Research Agents

Research progress increasingly comes from wrapping a fixed, often open-weight model in a generate-evaluate-revise loop against a machine-checkable objective, shifting competitive value from model weights to the search harness and the verifier.

forming · confidence 40 · Emerging (watchlist) · tracking since August 4, 2026 · updated August 4, 2026

Sign in to get threshold and movement alerts for this trend.

Why the conviction moved

  • Aug 4
    Strengthened +6

    OpenAI published machine-checkable Lean 4 formalizations for solutions to ten long-open mathematics problems from its Astra model, saying the underlying tokens would cost roughly two thousand dollars at its own API rates. A formal verifier removes the human referee from the loop entirely, which is the precondition for scaling generate-evaluate-revise search to arbitrary compute.

  • Aug 4
    Strengthened +7

    Tencent's Hyra agent reports beating the historical best on 29 of 55 open mathematics problems by running a generate, evaluate and revise loop on top of the company's existing open Hy3 model, with demonstration artifacts published on GitHub. The gain came from the harness rather than new weights, which is the direct claim of this thesis and now has a published artifact trail behind it.

  • Aug 4
    Strengthened +3

    AutoFOAM wraps a self-refinement loop around OpenFOAM, using the solver's own success or failure as the objective. It is the same harness-over-verifier pattern as the Lean and mathematics work, applied to an engineering simulator rather than a proof checker.

Source trail

  • Supporting · August 4, 2026

    A Self-Refining Agent Takes On OpenFOAM Configuration, One of Engineering's Reliable Time Sinks

    AutoFOAM wraps a self-refinement loop around OpenFOAM, using the solver's own success or failure as the objective. It is the same harness-over-verifier pattern as the Lean and mathematics work, applied to an engineering simulator rather than a proof checker.

    arXiv cs.AI
  • Supporting · August 4, 2026

    OpenAI Publishes Lean-Verified Proofs From Its Astra Model for Ten Long-Open Mathematics Problems

    OpenAI published machine-checkable Lean 4 formalizations for solutions to ten long-open mathematics problems from its Astra model, saying the underlying tokens would cost roughly two thousand dollars at its own API rates. A formal verifier removes the human referee from the loop entirely, which is the precondition for scaling generate-evaluate-revise search to arbitrary compute.

    AI Post (Telegram)

Unlock full source trail, score history, and daily updates.

1 more source in the full trail.

Unlock Trends

Affected regions & assets

Assets3 assetsUnlock Trends