acceptodds
Under review as a conference paper at ICLR 2027

Predicting IMP Masks at Initialization under Finite-Time Linearized Dynamics

Abstract

Iterative magnitude pruning (IMP) selects sparse masks from weights obtained after training, whereas pruning at initialization selects masks before weight training. We formulate this difference as a **finite-time terminal-ranking problem**: can the magnitudes used by IMP be determined from initialization-side quantities? Under fixed-Jacobian squared-loss gradient descent, the finite-time response specifies terminal weights from the initialization, data, learning rate, and training horizon. Applying the same magnitude-pruning rule recovers one-shot and iterative linearized IMP masks when the top- boundary is stable. We validate this correspondence in controlled and functional whole-network fixed-Jacobian models. On correlated-feature tasks, the predictor recovers teacher supports where the tested static initialization scores degrade, and finite-time ablations show that the converged solution can select different masks. Beyond the exact regime, initialization-side scores combined with architecture-based layer allocation produce competitive sparse retraining accuracy in several nonlinear LeNet and ResNet settings. Repeated IMP discoveries provide a reference for interpreting coordinate agreement under training randomness. A matrix-free implementation evaluates the predictor without explicit pseudoinverses. Together, these results connect posterior IMP selection to initialization-side trajectory prediction and distinguish exact mask recovery from functional transfer to nonlinear training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.