Predicting IMP Masks at Initialization under Finite-Time Linearized Dynamics
Abstract
Iterative magnitude pruning (IMP) selects sparse masks from weights obtained after training, whereas pruning at initialization selects masks before weight training. We formulate this difference as a **finite-time terminal-ranking problem**: can the magnitudes used by IMP be determined from initialization-side quantities? Under fixed-Jacobian squared-loss gradient descent, the finite-time response specifies terminal weights from the initialization, data, learning rate, and training horizon. Applying the same magnitude-pruning rule recovers one-shot and iterative linearized IMP masks when the top- boundary is stable. We validate this correspondence in controlled and functional whole-network fixed-Jacobian models. On correlated-feature tasks, the predictor recovers teacher supports where the tested static initialization scores degrade, and finite-time ablations show that the converged solution can select different masks. Beyond the exact regime, initialization-side scores combined with architecture-based layer allocation produce competitive sparse retraining accuracy in several nonlinear LeNet and ResNet settings. Repeated IMP discoveries provide a reference for interpreting coordinate agreement under training randomness. A matrix-free implementation evaluates the predictor without explicit pseudoinverses. Together, these results connect posterior IMP selection to initialization-side trajectory prediction and distinguish exact mask recovery from functional transfer to nonlinear training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.