acceptodds
Under review as a conference paper at ICLR 2027

Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws

Abstract

Transformers improve predictably with model scale, yet the internal mechanisms underlying this behavior remain unclear. We study Transformer residual-stream components through bias–diversity decomposition, viewing their projected outputs as an ensemble. This perspective characterizes how component accuracy, diversity, and importance jointly affect prediction quality. We further develop an information-theoretic formulation in which adding informative layers monotonically improves prediction-error bounds. We characterize departures from diminishing returns by a dependence-increase term and show that the bounds exhibit approximate diminishing marginal improvements when this term is small. This provides a layer-level mechanism qualitatively consistent with the diminishing returns observed in parameter scaling laws. Experiments across eight NLP tasks and multiple LLM families show systematic relationships among bias, diversity, and accuracy. We also find that component reweighting alters the bias–diversity balance and that recurrent layer sharing substantially reduces measured diversity. Overall, our results establish layer diversity as a useful lens for understanding how Transformer components cooperate as model depth increases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.