Post-Training's Diversity Tradeoff
As labs optimize models hard for reasoning and target behaviors, evidence recurs that post-training erodes behavioral diversity, forcing new objectives that preserve exploration for agentic use.
forming · confidence 40 · Emerging (watchlist) · tracking since July 23, 2026 · updated July 23, 2026
Why the conviction moved
- Jul 23Strengthened +5
A study finds reasoning fine-tuning collapses behavioral diversity in sequential game play, with supervised fine-tuning narrowing the range of moves models make. This directly instantiates the thesis that post-training for target behaviors erodes exploration needed for agentic use.
Source trail
Supporting · July 23, 2026
Study Finds Reasoning Fine-Tuning Collapses Behavioral Diversity in Game Play
A study finds reasoning fine-tuning collapses behavioral diversity in sequential game play, with supervised fine-tuning narrowing the range of moves models make. This directly instantiates the thesis that post-training for target behaviors erodes exploration needed for agentic use.
arXiv cs.CL
Unlock full source trail, score history, and daily updates.
Unlock Trends