# Researchers Argue Agent Safety Belongs at Runtime, Not in the Training Run

A new paper contends that reinforcement learning from human feedback, direct preference optimization and Constitutional AI are structurally insufficient for agents that execute code and mutate files, and a separate wave of tools is already moving control outside the model.

- Published: 2026-08-13T06:26:09.114Z
- Canonical: https://polylog.news/ai/2026-08-13/researchers-argue-agent-safety-belongs-at-runtime-not-in-the
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CR](https://arxiv.org/abs/2608.11274), [Polylog editors](https://polylog.news)

A paper posted to arXiv, Agent Safety Should Be a Runtime Contract, argues directly against the dominant approach. Safety instilled during training through reinforcement learning from human feedback (RLHF), direct preference optimization (D…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-13/researchers-argue-agent-safety-belongs-at-runtime-not-in-the (subscription information: https://polylog.news/pricing).