Morning Edition · Monday, August 24, 2026Published at 3:20 AM EDT · New York
The aiXamine authors give an example of a model scoring 99.3 on safety alignment while refusing one in three benign requests, and a second paper argues refusal training leaves harmful knowledge intact underneath.

The critical failure modes of deployed large language models are cross-dimensional, argue the authors of aiXamine, a unified black-box evaluation framework posted to arXiv on 24 August. Their illustrative case is blunt: a model can score 99…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
Oversight and Evaluation Lag Accelerating AI Capabilities
Over the next 3-6 months, evidence mounts that governance, evaluation, and agent-safety methods are failing to keep pace with capability growth, driving investment in interpretability, agent-manipulation benchmarks, and institutional-reform proposals.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.