# Amodei Commits Anthropic to Embedded Outside Evaluators, and OpenAI Says It Will Match

The essay reverses Amodei's 2023 position against slowing down and comes five days after a pretraining researcher resigned from Anthropic saying the industry is not in control of what it is building.

- Published: 2026-09-14T06:22:55.303Z
- Canonical: https://polylog.news/ai/2026-09-14/amodei-commits-anthropic-to-embedded-outside-evaluators-and
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news)

Dario Amodei, the chief executive of Anthropic, published an essay on September 12 titled [We Must Pace the Frontier](https://darioamodei.com/post/we-must-pace-the-frontier), arguing that the AI industry should deliberately slow the pace at which it advances model capability. He set out a three-part plan and committed Anthropic unilaterally to the first part: giving third-party evaluators permanent, employee-level access so they can verify that the company follows its own safety commitments, report incidents, and assess model alignment during training rather than only after a model is finished.

The Russian-language technical channel AI ML Big Data [summarized the reversal plainly](https://t.me/ai_machinelearning_big_data/10917): in 2023 Amodei considered slowing down premature, and he now considers it necessary. The essay identifies three risks that this pacing is meant to guard against, the loss of control over AI systems, misuse for cyberattacks and biological weapons, and severe economic disruption.

The timing matters more than the argument. Jacob Coxon, who spent roughly three years on pretraining research at OpenAI and then Anthropic, [resigned on September 8](https://time.com/article/2026/09/09/ai-anthropic-openai-jacob-coxon/) and posted a seven-part statement on X saying neither company is acting responsibly. Axios [reported](https://www.axios.com/2026/09/09/anthropic-researcher-ai-warning-interview) that he surrendered his equity to leave. Amodei addressed the post directly on September 13, saying [he agrees with Coxon much more than he disagrees](https://t.me/aipost/8137) and that Coxon was describing an industry-wide race dynamic rather than accusing Anthropic specifically.

Reaction split within hours. Elon Musk endorsed the essay, and Sam Altman, the chief executive of OpenAI, committed OpenAI to match the embedded-evaluator program. Critics including Emad Mostaque, the founder of Stability AI, called the proposal well-intentioned but not enforceable, since nothing in it binds a lab that declines to participate. Altman, asked in an interview why the major lab heads do not simply meet to settle AI safety, [said such a meeting "will happen"](https://t.me/aipost/8139) but declined to pre-announce private discussions. What is verified today is one company's access commitment and one rival's public agreement to match it. What is not verified is any mechanism for a third lab, a Chinese open-weight developer, or a well-funded startup to be held to the same standard.

## What this means

Embedded third-party evaluators change the cost structure of frontier training, not just how the effort is perceived publicly. If OpenAI and Anthropic both seat outside reviewers inside training runs, the two largest closed labs absorb schedule risk and disclosure obligations that open-weight developers and smaller rivals do not carry, which widens the release-cadence gap in favor of anyone outside the arrangement. Two outcomes are possible: either evaluator access becomes a de facto licensing condition that regulators later codify, raising the barrier to entry and protecting incumbent margins, or it stays voluntary and unenforced, in which case it serves only to protect participants' reputations while capability release schedules continue unchanged. Which outcome occurs will be visible in whether any government cites the commitment in rulemaking.

## What to watch

- Whether OpenAI publishes the actual terms of its matching commitment, including which evaluators get access and whether they can halt a training run, since a commitment without a stop authority is a reporting arrangement rather than a control.
- Whether Google DeepMind, Meta Superintelligence Labs, or any Chinese lab joins, because a two-lab arrangement sets no industry floor and shifts release timing advantage to non-participants.
- Whether Anthropic's next frontier release slips relative to its recent cadence, which would be the first measurable evidence that pacing has a real cost rather than a rhetorical one.
