# Post-Training's Diversity Tradeoff

As labs optimize models hard for reasoning and target behaviors, evidence recurs that post-training erodes behavioral diversity, forcing new objectives that preserve exploration for agentic use.

- Conviction: 40 / 100 (forming)
- Horizon: Emerging (watchlist)
- Tracking since: 2026-07-23T00:00:00.000Z
- Last updated: 2026-07-23T05:54:22.032Z
- Canonical: https://polylog.news/ai/trends/rl-posttraining-diversity-tradeoff
- Publisher: Polylog
- Affected regions: Global

## Recent evidence

- [confirms] Study Finds Reasoning Fine-Tuning Collapses Behavioral Diversity in Game Play (2026-07-23): A study finds reasoning fine-tuning collapses behavioral diversity in sequential game play, with supervised fine-tuning narrowing the range of moves models make. This directly instantiates the thesis that post-training for target behaviors erodes exploration needed for agentic use.
