# A New Benchmark Measures Whether Frontier Models Know They Are Being Tested

EvalDetectBench, from LASR Labs and the UK AI Security Institute, scores two things at once: how reliably a model detects an evaluation, and how detectable each benchmark is.

- Published: 2026-09-03T06:26:17.363Z
- Canonical: https://polylog.news/ai/2026-09-03/a-new-benchmark-measures-whether-frontier-models-know-they-a
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.AI](https://arxiv.org/abs/2609.01611), [Anthropic Research](https://www.anthropic.com/research/team/frontier-red-team)

Researchers at LASR Labs, the University of Pennsylvania and the UK AI Security Institute released EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness in frontier large language models. Evaluation awareness is…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-09-03/a-new-benchmark-measures-whether-frontier-models-know-they-a (subscription information: https://polylog.news/pricing).