# A New Benchmark Tests Whether Chemistry Agents Can Be Steered Into Harm Over Long Workflows

The preprint argues the relevant safety question has moved from whether a model answers a dangerous question to whether a multi-step discovery pipeline can be steered toward a dangerous output.

- Published: 2026-09-14T06:22:55.303Z
- Canonical: https://polylog.news/ai/2026-09-14/a-new-benchmark-tests-whether-chemistry-agents-can-be-steere
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv](https://arxiv.org/abs/2609.11952)

A preprint posted on September 14 introduces ChemMat-AgentSafetyBench, an evaluation of long-horizon attacks and defenses against agents operating in chemistry and materials science. Its framing is the significant part. Chemistry and materi…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-09-14/a-new-benchmark-tests-whether-chemistry-agents-can-be-steere (subscription information: https://polylog.news/pricing).