Only the Boundary Matters: Efficient Mask Learning for Sparse Neural Networks
Abstract
Sparse neural networks construct subnetworks by selectively retaining weight connections and are widely used for model compression, inference acceleration, and parameter-efficient adaptation. Binary masks implement connection selection in sparse networks, while learned masks use continuous scores to adapt this selection to data or tasks. However, discrete masks still rely on dense floating-point scores, gradients, and optimizer states, whose update overhead grows with the number of candidate weights. We identify a simple rank-stability phenomenon: under bounded score updates, score intervals that cannot cross the Top- threshold correspond to fixed mask decisions, and only boundary scores can change their inclusion status. Based on this observation, we propose Boundary, an active-set optimization framework for learned binary masks. It stores determined decisions as fixed bits, maintains continuous optimization state only for boundary scores, and completes selection using the remaining quota; periodically committing state and rebuilding the active set allows previously fixed coordinates to re-enter optimization. Experiments on ImageNet and random-weight subnetworks show that Boundary substantially reduces controller state while maintaining accuracy close to the original mask optimization. On Pythia-1.4B, its language-model quality is close to Full-score; in an independent short-context timing benchmark, controller state and peak allocated memory decrease by 80.6% and 67.1%, respectively, with a measured training-loop speedup of 9.20. Boundary turns Top- rank stability into a concrete rule for allocating optimization resources according to whether mask decisions can change, providing a direct way to reduce mask-training cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.