Morning Edition · Wednesday, July 15, 2026Published at 1:45 AM EDT · New York
Under a token-controlled protocol, wrappers that differ only in formatting shift model scores enough to reorder benchmarks, and the authors propose a Format Sensitivity Index to measure it.
A paper on the Format Sensitivity Index documents a problem that undermines confidence in public leaderboards. Prompt wrappers that differ only in formatting, not in content, can change a model's score enough to reverse the ranking conclusi…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
Oversight and Evaluation Lag Accelerating AI Capabilities
Over the next 3-6 months, evidence mounts that governance, evaluation, and agent-safety methods are failing to keep pace with capability growth, driving investment in interpretability, agent-manipulation benchmarks, and institutional-reform proposals.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.