Understanding Network Size Scaling through Linear Region Geometry
Abstract
Neural network test losses are observed to evolve as power laws with the network size, but there is no theory that can fully predict their parameters after training. We present an empirical study of the learned properties that describes the scaling by exploiting the linear region partition of ReLU networks. We show that trained neural networks do not constitute ideal piecewise-linear approximators and identify that the regression loss is dominated by deviations of regional affine offsets from their optimal values. This offset error is described locally as the first-order variation of the target over a characteristic regional scale, and its mean depends approximately on the square of that same scale averaged over high-gradient regions. These results connect neural scaling to measurable properties of the piecewise-affine structure of networks and target geometry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.