Rethinking Inactivity in Residual Networks
Abstract
Controlling activation variance does not determine how many features are zero on every input. Such features contribute no signal to a linear readout on that input set. We develop an inverse theory that chooses a residual initialization scale from a prescribed fraction of inactive features. For Gaussian branches followed by a terminal ReLU, we prove convergence of the joint empirical feature law on any fixed input set as state width and depth grow at arbitrary relative rates, allowing branch width to remain finite. The law yields unique inverses for marginal inactivity on positive inputs and, for ReLU or leaky ReLU branches, joint inactivity on positive or centered Gaussian input laws. The marginal inverse explains why input rescaling requires opposite corrections for GELU and SELU, while ReLU and leaky ReLU are invariant. We also construct positive non-Gaussian input laws with identical pairwise terminal feature distributions but different unused-feature fractions and expected readout ranks at every fixed budget of sampled limiting features. Experiments verify calibration and the predicted fitting differences. The results identify the input and activation information needed to control initial feature usage beyond variance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.