# Researchers Propose Auditing AI Scientists by Their Process Traces, Not Their Papers

OpenDiscoveryTrace argues that benchmarks scoring only final code, hypotheses or write-ups make scientific claims impossible to audit, as national laboratories put vision models into live research pipelines.

- Published: 2026-09-10T06:24:02.019Z
- Canonical: https://polylog.news/ai/2026-09-10/researchers-propose-auditing-ai-scientists-by-their-process
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.AI](https://arxiv.org/abs/2609.09203), [arXiv cs.CL](https://arxiv.org/abs/2609.09264), [Meta AI](https://ai.meta.com/blog/genesis-mission-lawrence-berkeley-national-laboratory-segment-anything-dino/)

Existing benchmarks for autonomous AI scientists grade the output and discard everything else. OpenDiscoveryTrace makes the case that this is the wrong unit of evaluation, because a correct hypothesis reached through a flawed or fabricated…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-09-10/researchers-propose-auditing-ai-scientists-by-their-process (subscription information: https://polylog.news/pricing).