# New Papers Push Back on Paying Frontier Prices to Grade Model Output

One study asks whether cheap open-weight models can judge natural-language mathematical proofs reliably, as forecasters put the model evaluation tools market near $1.15 billion in 2025.

- Published: 2026-08-04T06:16:46.804Z
- Canonical: https://polylog.news/ai/2026-08-04/new-papers-push-back-on-paying-frontier-prices-to-grade-mode
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2608.00004), [arXiv cs.CL](https://arxiv.org/abs/2608.00005), [Polylog editors](https://polylog.news)

Grading is now a recurring cost in evaluating reasoning systems, and frontier judges are expensive. A paper posted to arXiv on August 4 asks the direct question: can cheap open-weight models serve as reliable judges of natural-language math…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-04/new-papers-push-back-on-paying-frontier-prices-to-grade-mode (subscription information: https://polylog.news/pricing).