Pruning Close to Home: Distance from Initialization impacts Lottery Tickets
Abstract
The Lottery Ticket Hypothesis states that there exist sparse subnetworks (called 'winning' Lottery Tickets) within dense randomly initialized networks that, when trained under the same regime, achieve similar or better validation accuracy as the dense network. It has been shown that for larger networks and more complex datasets, these Lottery Tickets cannot be found in random initializations, but that they require lightly pretrained weights. In this paper, we aim to circumvent this pretraining phase by exploring techniques to improve the quality of the pruning masks that emerge during the Mask Search phase. We start by demonstrating that a simple choice of hyperparameters during the search can significantly impact the trainability of the sparse network, and that the trends that hold for searching are contradictory with the best hyperparameters for training. This leads us to explore learning distance as an explanation for this, which is influenced by hyperparameter selection, and demonstrate its impact via explicit regularization. We find that a lower learning distance in combination with limited sign flips during pruning results in sparse networks that are better trainable, and that after several pruning rounds there even emerge useful features as-is in the untrained sparse network. Finally, we combine these observations to significantly reduce the extraction cost of a highly sparse network.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.