# A Safety Evaluation Paper Argues Single-Dimension Scores Hide How Aligned Models Actually Fail

The aiXamine authors give an example of a model scoring 99.3 on safety alignment while refusing one in three benign requests, and a second paper argues refusal training leaves harmful knowledge intact underneath.

- Published: 2026-08-24T07:20:23.472Z
- Canonical: https://polylog.news/ai/2026-08-24/a-safety-evaluation-paper-argues-single-dimension-scores-hid
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CR (aiXamine)](https://arxiv.org/abs/2608.20554), [arXiv cs.AI (Truth Lies Deep)](https://arxiv.org/abs/2608.20378), [arXiv (prior aiXamine version)](https://arxiv.org/abs/2504.14985), [Anthropic](https://www.anthropic.com/news/hard-questions)

The critical failure modes of deployed large language models are cross-dimensional, argue the authors of aiXamine, a unified black-box evaluation framework posted to arXiv on 24 August. Their illustrative case is blunt: a model can score 99…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-24/a-safety-evaluation-paper-argues-single-dimension-scores-hid (subscription information: https://polylog.news/pricing).