WHERE DOES OVERFITTING LIVE? DEPTHD-EPENDENT STRUCTURE IN PRUNED NEURAL NETWORKS
Abstract
Although structured pruning is known to improve generalization in overparameterized networks, which parameters carry the overfitting burden—and why pruning consistently finds them—remains unclear. We show that, during training, the deepest active block undergoes a simultaneous collapse in both BN scale factors and channel L2-norms that is absent in shallower blocks: overfittingrelated capacity concentrates toward the network’s deepest stage, which we term the noise sink. On ResNet-18, all three criteria exclusively target the deepest stage until the architecturally determined 53.33% threshold; on ResNet-34, exclusivity is criterion-dependent; on VGG-16, the two deepest stages are cotargeted from the first cuts. Crossing the stage-exhaustion boundary triggers a sharp, criterion-dependent generalization transition. Sequential block removal further reveals that the noise sink role is not fixed to any specific block: after the deepest block is removed, the next deepest undergoes the same dual collapse and becomes the new pruning target, confirmed on ResNet-18, ResNet-34, and VGG-16. Controlled experiments show that Block 4 suppression strengthens monotonically with training data volume, the L2 component persists under SGD+momentum, and the pattern is invariant to label noise—indicating a general property of over-parameterized network convergence rather than a small-data or optimizer artifact. Code is available at https://anonymous.4open. science/r/depth-structure-pruning-CB93/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.