Supervision Density Shapes Transformer Solutions and Long-Context Generalization
Abstract
Which tokens receive supervision is typically treated as an efficiency choice: supervising fewer tokens can reduce training cost while preserving performance. Here, we show that supervision density also determines how a Transformer learns to solve a reasoning task, with direct consequences for long-context generalization. We train 12-layer Transformers on a controlled multi-hop variable-binding task while varying the fraction of suffix tokens contributing to the loss, from standard all-token prediction to last-token supervision, and evaluate generalization to both longer contexts and deeper reasoning chains. We uncover a strikingly non-monotonic relationship between supervision, learning speed, and generalization. Intermediate supervision densities converge substantially faster than either all-token or answer-only training, yet the fastest-converging model is not the best generalizer. At convergence, sparse 25% supervision yields the strongest long-context generalization, whereas prolonged training eventually makes the last-token model the strongest generalizer despite validation accuracy having saturated much earlier. Mechanistic analyses reveal that these learning regimes steer models toward qualitatively different solutions: suffix-supervised models predominantly write the answer through the MLP pathway, whereas last-token supervision produces an attention-dominated solution. Crucially, attention remains causally necessary across all regimes, and the strongest generalizer at convergence is the model that most strongly requires both attention and MLP pathways. Together, our results show that supervision density does more than control training efficiency: it selects among memory- and structure-based solutions, thereby shaping when and how Transformers generalize beyond their training distribution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.