Harmonic Variance Embedding: A Positional-Style Representation for Counteracting Depth Dilution in Pre-Norm Transformers
Abstract
Pre-norm Transformers have become the backbone of modern large language models, yet their residual stream is not normalized, causing activation variance to grow with depth and gradually diminishing the contribution of each sublayer—a phenomenon known as the curse of depth. In this work, we propose Harmonic Variance Embedding (HVE), a simple and lightweight method that modulates residual dynamics based on input statistics. HVE computes the per-token variance of the residual stream, encodes it using a positional-encoding-style embedding mechanism.HVE maps the resulting representation into the hidden states of intermediate layers via a small MLP with zero-initialized output so that each layer can recognize the variance structure of its own input and adaptively choose the variance of the output it generates , allocating recognizable bandwidth by importance × scale in a matched way so as to partly solve the curse of depth. The positional-style representation is bounded and smooth, effectively suppressing outlier sensitivity while preserving a structured feature of variance magnitude. HVE does not alter the core architecture and adds only a negligible amount of parameters. It importantly mitigates the severe generalization degradation caused by naive variance injection. These findings suggest that our positional-encoding-style encoding of activation statistics offers a practical and robust alternative to explicit normalization or learned scaling mechanisms for deep Pre-norm Transformers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.