# Provenance Becomes Training Infrastructure

As machine-generated text takes over a growing share of the open web, watermarking and provenance tracking move from optional safety features into required components of training pipelines, and clean human data becomes a priced, contested input.

- Conviction: 41 / 100 (weakening)
- Horizon: Emerging (watchlist)
- Tracking since: 2026-08-25T00:00:00.000Z
- Last updated: 2026-08-27T14:00:31.454Z
- Canonical: https://polylog.news/ai/trends/synthetic-data-provenance
- Publisher: Polylog
- Affected regions: Global

## Recent score history

- 2026-08-26: 43
- 2026-08-27: 41

## Recent evidence

- [confirms] OpenAI Bans Russia-Origin Accounts Running a Fake Israeli Think Tank and a Pro-Russia 'Sovereignty Index' (2026-08-26): The Russia-origin operators OpenAI banned instructed ChatGPT to remove signs that the output was machine-generated before publishing it as think-tank research. Deliberate laundering of machine text into the open web is the failure mode that makes provider-side watermarking and provenance, rather than post-hoc detection, the only workable control for training-corpus hygiene.
- [confirms] A New Survey Catalogues Model Collapse as Labs Formalize Content Provenance (2026-08-25): A new survey catalogues model collapse — what happens when generative models train on AI-synthesized data — landing the same week Anthropic published how it watermarks Claude's text output. The failure mode and the provenance mechanism are being formalized together, which is what turns watermarking from a safety nicety into a data-pipeline filter labs need for their own training runs.
