# Study Finds LLM Judges Disagree With Themselves on Repeated Identical Runs

Re-running the same evaluation many times exposes run-to-run instability in the LLM-as-a-Judge method that underpins leaderboards and reward models.

- Published: 2026-06-15T07:00:34.492Z
- Canonical: https://polylog.news/ai/2026-06-15/study-finds-llm-judges-disagree-with-themselves-on-repeated
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2606.13685)

A new paper, "The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation," studies what happens when the same model is asked to judge the same outputs repeatedly, in work posted to arXiv. The authors run repeated identical evalu…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-06-15/study-finds-llm-judges-disagree-with-themselves-on-repeated (subscription information: https://polylog.news/pricing).