Fisher Geometry of Activations for Deep Representation Learning: Diagnosis and Mechanism
Abstract
Residual connections and normalization are central to deep-network training, yet their effects on layer-wise activation spectra remain incompletely characterized. The primary contribution of this paper is a staged diagnosis of those spectra and a local account of the measured patterns. We measure effective rank, participation ratio, and condition number of activation covariance, together with the Fisher information of a local Gaussian activation model. Under that Gaussian model, Fisher information for the mean is the precision, and layer linear Fisher information (LFI) is the squared length of a mean shift in this metric. LFI is separate from parameter Fisher information. Covariance eigenvalues are reciprocals of Fisher eigenvalues. On the tested MLP, CNN, and attention encoders, plain stacks often concentrate covariance energy, residual variants often have broader spectra, and per-channel standardization changes diversity and conditioning. These are coordinate-dependent measurements: normalization rescales those coordinates, so broader covariance is not invariant evidence of preserved information. Locally, a residual identity path changes gain propagation, but the leading Jacobian singular value does not determine effective rank. FisherNorm learns a per-channel standardization exponent from the task loss. Under the shared placement protocol, it is comparable to the strongest baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.