acceptodds
Under review as a conference paper at ICLR 2027

Normalization as Gauge Fixing: A Variational Characterization of Batch Normalization

Abstract

Normalization layers appear in nearly every deep network, yet why they help is still unsettled: BN's original internal-covariate-shift explanation has been challenged, and the six later accounts each rest on measurements, yet none is load-bearing alone, since networks train competitively without a normalizer. Recent work has pinned down the geometry of what the operator produces, the centered sphere its outputs lie on, but not the problem that geometry solves. We supply that problem. Rescaling a neuron's incoming weights leaves a normalized network's function unchanged, so many parameter settings compute the same map and normalization picks one: gauge fixing, in the language of constrained mechanics. For one batch and one channel, batch normalization's two steps are the exact closed-form solution of a strictly convex selection problem over this redundancy, verified symbolically and numerically. Short proofs follow: why a zero-variance channel breaks the operator, what epsilon repairs, why depth needs no iteration, and when a centering folds into the next layer for free. The objective also says why this operator and not a near neighbour: other divisors remove the same redundancy and leave the forward pass unchanged, but only the standard deviation has a backward pass that distorts no direction, and the alternatives lose accuracy in proportion to their distortion, up to points. The same structure splits the benefit into four parts, one of which we measure cleanly: holding the forward pass identical and changing only the backward pass, the batch-coupled gradient terms are worth to accuracy points on CIFAR-10.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.