# Training Objectives Move Beyond Next-Token Prediction

Architectures that predict in latent or concept space alongside next-token training keep posting reasoning gains at fixed parameter counts, so expect recurring non-token-level objectives to scale further and for the pretraining objective — not data volume or parameter count — to become a contested axis of frontier model design.

- Conviction: 38 / 100 (weakening)
- Horizon: Emerging (watchlist)
- Tracking since: 2026-09-11T00:00:00.000Z
- Last updated: 2026-09-14T14:04:09.672Z
- Canonical: https://polylog.news/ai/trends/beyond-next-token-training-objectives
- Publisher: Polylog
- Affected regions: Global

## Recent score history

- 2026-09-13: 39
- 2026-09-14: 38

## Recent evidence

- [confirms] A Collapse-Free Video Encoder Matches V-JEPA 2 With Up to 20.8 Times Less Pretraining Compute (2026-09-13): LeVJEPA replaces the usual self-supervised heuristics in latent-space video prediction with a single regularizer that provably rules out representation collapse. Removing the main theoretical objection to joint-embedding predictive objectives makes the training objective, rather than data or parameter count, the lever producing the gain.
- [confirms] A Latent-Space Language Model Scales to 8.9 Billion Parameters and Beats OLMo-3 on Reasoning (2026-09-11): NCP-ArchPreview scales latent-space concept-level prediction to 8.9 billion parameters, beating OLMo-3 on reasoning benchmarks including a 5.99-point GSM8K improvement, the largest published run of this objective class to date.
