Morning Edition · Thursday, August 13, 2026Published at 2:26 AM EDT · New York
A new paper contends that reinforcement learning from human feedback, direct preference optimization and Constitutional AI are structurally insufficient for agents that execute code and mutate files, and a separate wave of tools is already moving control outside the model.

A paper posted to arXiv, Agent Safety Should Be a Runtime Contract, argues directly against the dominant approach. Safety instilled during training through reinforcement learning from human feedback (RLHF), direct preference optimization (D…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
Oversight and Evaluation Lag Accelerating AI Capabilities
Over the next 3-6 months, evidence mounts that governance, evaluation, and agent-safety methods are failing to keep pace with capability growth, driving investment in interpretability, agent-manipulation benchmarks, and institutional-reform proposals.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.