Understanding the Predictive-State Bottlenecks of Transformer Models
Abstract
Sequence-prediction tasks typically require compressing an observation history into a compact internal state, making the state's dimensionality and organization central to neural model capacity. We study whether the residual width required for accurate prediction is governed by the latent-state count, the full posterior dimension, or the intrinsic dimension of the predictive state, using Transformers trained on hidden Markov processes. Across controlled families that separate these quantities, the width transition tracks predictive-state rank. Exact filtering constructions and bottleneck results connect this transition to the geometry of predictive representations, while readout and architectural controls distinguish state capacity from output parameterization. Also, we observe that the training is spectrally ordered, and a probe-calibrated analysis reveals a width-dependent mechanism: successive Transformer blocks increasingly favor high-variance belief directions; below rank, this becomes suppression of low-variance modes beyond an equal-treatment probe null, while sufficiently wide, adequately fitted models reverse the bias. Causal interventions separately show that the represented belief subspace contributes to prediction. In the tested contractive regime, filtering is strongly width-sensitive and only weakly depth-sensitive.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.