# A New Benchmark Finds Frontier Models Fail About One in Three Research-Integrity Decisions Under Pressure

IntegrityBench tested 18 model variants across 36 paired tasks with escalating institutional pressure, and reports that neither model scale nor reasoning ability reliably fixed the failures.

- Published: 2026-08-15T06:14:29.474Z
- Canonical: https://polylog.news/ai/2026-08-15/a-new-benchmark-finds-frontier-models-fail-about-one-in-thre
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.AI](https://arxiv.org/abs/2608.12345)

A paper posted to arXiv, Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists, introduces IntegrityBench, a benchmark that tests whether language models uphold research integrity norms when circumstances push again…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-15/a-new-benchmark-finds-frontier-models-fail-about-one-in-thre (subscription information: https://polylog.news/pricing).