Caffeinate: Recovering Dormant Subnetworks in Modern Convolutional Neural Networks
Abstract
One-shot pruning is understood to degrade accuracy exclusively due to the removal of critical weights from a model. Evidence from batch-normalized CNNs supports a different account. The pruned model retains a high-performing sub-network that signal collapse, which is a progressive, depth-wise decay of activation variance, renders dormant, and recalibrating the BatchNorm statistics recovers it. Modern convolutional architectures, such as ConvNeXt, FocalNet, HorNet, and RDNet, adopt LayerNorm instead of BatchNorm, which computes statistics per sample at inference time. LayerNorm contains no stored statistics to recalibrate. We demonstrate that signal collapse exists in these networks and that damage due to pruning lies in the per-channel output moments, the subspace the LayerNorm affine parameters control. We develop Caffeinate, a closed-form correction of LayerNorm's affine parameters, which is applied once to a pruned network to restore the original per-channel output moments from only forward passes, without gradients or labels. Since Caffeinate rewrites only the affine parameters, it can be applied to any one-shot pruning mask. We validate this on magnitude pruning (MP) and CHITA, a state-of-the-art impact-based pruning method which uses second order information to select which parameters to prune and to update the unpruned parameters. Caffeinate recovers up to accuracy over MP across twelve ImageNet networks. MP corrected by Caffeinate outperforms CHITA on ten of twelve models at and all twelve models at , by up to accuracy, with a memory footprint of – GB against CHITA's – GB. When applied to CHITA, Caffeinate improves the accuracy of the pruned model by up to . This demonstrates that dormant high-performing sub-networks exist after one-shot pruning, regardless of the pruning method, and that correcting per-channel output moments can recover accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.