Normalization Couples What the Objective Separates: Representation Selection in Group Composition Networks
Abstract
We study when the representation selected by a normalized network can be predicted by scoring candidate representations independently. The model is a two-layer quadratic network on a finite group's multiplication table, with real Fourier sectors as the candidates. At small logit scale the leading objective has no cross-sector terms, but the sphere-projection multiplier depends on the sum of all sector terms. A sector-local rule may score each candidate from its own Fourier coefficients and from every sector's per-layer magnitudes, but not from another sector's signs, phases or matrix directions. Let the joint input sphere have metric weight and the output sphere weight one; ordinary normalization has . For groups with at least three nontrivial real Fourier sectors, sector-local rules recover the leading selector almost everywhere if and only if , the degree-matched weight. At every other fixed weight, the errors of measurable sector-local rules under uniform initialization are bounded away from zero. The proof reflects the output block of a third sector that wins at neither start and thereby switches the winner between two rivals whose permitted inputs are unchanged. Under ordinary normalization, for each fixed group and finite width and almost every initialization, exact cross entropy inherits the leading labels at its infinite-time endpoints as the scale vanishes. At width one, with three distinct nontrivial real linear characters, the optimal sector-local error of continuous exact cross entropy is positive at every sufficiently small positive scale, even at , where it converges to zero with the scale, and uniformly over any fixed compact set of weights. On two-bit XOR under ordinary normalization, a simple sector-local rule errs on 1.7% of sampled initializations, while the leading selectors of jointly and separately normalized networks have identical label frequencies but disagree on an estimated 18% of paired starts. These examples separate independent representation competition from both approximate prediction and agreement of label frequencies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.