Polylog
The Polylog AI Intelligence Brief

Morning Edition · Thursday, July 23, 2026Published at 2:03 AM EDT · New York

Study Finds Supervised Fine-Tuning Collapses Behavioral Diversity in Sequential Decisions

In controlled game-play suites, fine-tuned models narrow toward a few moves, a failure mode that reasoning traces can worsen rather than fix.

Study Finds Supervised Fine-Tuning Collapses Behavioral Diversity in Sequential Decisions

A new paper, When Reasoning Narrows the Move, studies how supervised fine-tuning (SFT) affects behavioral diversity in sequential decision-making. Using a controlled suite of games, the authors document what they call diversity collapse: af…

Continue the AI Intelligence Brief

Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.

  • 5 AI intelligence signals a day
  • Frontier labs, compute, and chips
  • Model releases and AI infrastructure
  • Source-grounded analysis with confidence labels

The Global Intelligence Brief stays free.

Part of a tracked trend

Post-Training's Diversity Tradeoff

As labs optimize models hard for reasoning and target behaviors, evidence recurs that post-training erodes behavioral diversity, forcing new objectives that preserve exploration for agentic use.