# New Benchmarks Argue Enterprise AI Evaluations Are Measuring the Wrong Thing

One paper shows text-to-SQL systems scoring above 89 percent on academic benchmarks face untested enterprise dialects and produce wrong answers that trigger no error, while another finds memory evaluations ignore how evidence is presented to the model.

- Published: 2026-08-26T06:22:56.640Z
- Canonical: https://polylog.news/ai/2026-08-26/new-benchmarks-argue-enterprise-ai-evaluations-are-measuring
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.AI](https://arxiv.org/abs/2608.23569), [arXiv cs.AI](https://arxiv.org/abs/2608.23568), [arXiv cs.LG](https://arxiv.org/abs/2608.23660)

Three papers posted to arXiv on Wednesday approach the same problem from different directions: the evaluations enterprises rely on to justify deployment do not measure the conditions those systems will actually meet. ESQ-Bench makes the sha…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-26/new-benchmarks-argue-enterprise-ai-evaluations-are-measuring (subscription information: https://polylog.news/pricing).