Polylog
← Trends

Post-Training's Diversity Tradeoff

As labs optimize models hard for reasoning and target behaviors, evidence recurs that post-training erodes behavioral diversity, forcing new objectives that preserve exploration for agentic use.

forming · confidence 40 · Emerging (watchlist) · tracking since July 23, 2026 · updated July 23, 2026

Why the conviction moved

  • Jul 23
    Strengthened +5

    A study finds reasoning fine-tuning collapses behavioral diversity in sequential game play, with supervised fine-tuning narrowing the range of moves models make. This directly instantiates the thesis that post-training for target behaviors erodes exploration needed for agentic use.

Source trail

  • Supporting · July 23, 2026

    Study Finds Reasoning Fine-Tuning Collapses Behavioral Diversity in Game Play

    A study finds reasoning fine-tuning collapses behavioral diversity in sequential game play, with supervised fine-tuning narrowing the range of moves models make. This directly instantiates the thesis that post-training for target behaviors erodes exploration needed for agentic use.

    arXiv cs.CL

Unlock full source trail, score history, and daily updates.

Unlock Trends

Affected regions & assets

RegionsGlobal