acceptodds
Under review as a conference paper at ICLR 2027

Tracing Information Processing in Transformers with Stochastic Normalization

Abstract

Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied to downstream function: nearby states may produce different behaviors, while distant states may behave similarly. We instead give representations volume, turning similarity into statistical distinguishability. Overlapping stochastic representations necessarily induce overlapping downstream distributions, grounding latent comparison in model function and bringing it under information-theoretic tools such as the data-processing inequality. We realize this idea in pretrained transformers through a light-touch modification to normalization: at each residual-stream read, normalize the state, add isotropic Gaussian noise, and renormalize. During distillation fine-tuning, one learned allocation parameter per read distributes a fixed global rate budget across the processing stack. The resulting model can be viewed as transformer blocks reading the residual stream with learned finite precision. This stochastic construction supports complementary views of transformer computation: tracing which input distinctions remain accessible to downstream components, and tracking how variability introduced at different locations propagates and mixes through the network. Across vision and language transformers, these views reveal structured, component-specific patterns of information preservation and processing. Together, our results establish finite-precision stochastic representations as a functionally grounded lens on transformer computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.