Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
Abstract
Transformers improve predictably with model scale, yet the internal mechanisms underlying this behavior remain unclear. We study Transformer residual-stream components through bias–diversity decomposition, viewing their projected outputs as an ensemble. This perspective characterizes how component accuracy, diversity, and importance jointly affect prediction quality. We further develop an information-theoretic formulation in which adding informative layers monotonically improves prediction-error bounds. We characterize departures from diminishing returns by a dependence-increase term and show that the bounds exhibit approximate diminishing marginal improvements when this term is small. This provides a layer-level mechanism qualitatively consistent with the diminishing returns observed in parameter scaling laws. Experiments across eight NLP tasks and multiple LLM families show systematic relationships among bias, diversity, and accuracy. We also find that component reweighting alters the bias–diversity balance and that recurrent layer sharing substantially reduces measured diversity. Overall, our results establish layer diversity as a useful lens for understanding how Transformer components cooperate as model depth increases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.