acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Attend to: Training Dynamics of Multi-Token Prediction

Abstract

Multi-token prediction (MTP) has improved next-token prediction in several large-scale settings, but how its additional supervision changes contextual cue selection remains unclear. In this work, we study a one-layer linear-attention model with two prediction heads sharing a contextual representation. We analyze a two-stage procedure that learns token-to-future associations in the readouts, then freezes the readouts and trains shared attention by gradient flow. Under explicit structural assumptions, we characterize the max-margin solutions governing attention-score growth. We show that MTP assigns leading logarithmic growth only to cues jointly predictive of both future tokens. In contrast, next-token prediction (NTP) assigns leading coefficients according to the occurrence and overlap of next-token compatible cues, allowing ambiguous cues to receive as much or more leading growth as jointly predictive ones. We derive the resulting prediction rules on recombined contexts. An illustrative example shows that, as training time grows, the correct next-token probability tends to one under MTP and to zero under NTP on the same context, even when inference uses only MTP's first head. Synthetic experiments support the predicted cue preferences, max-margin directions, and prediction separation, while TinyStories experiments illustrate learned token-to-future associations and attention differences between NTP and MTP.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.