acceptodds
Under review as a conference paper at ICLR 2027

Sparsity without Sweeps: Adaptive Regularization for Mirror Descent

Abstract

Sparse training reduces the memory and computational costs of deep neural networks. However, sparse optimization methods, e.g., those adding an penalty, often control sparsity only indirectly through a regularization parameter , whose mapping to the final sparsity rate is highly non-trivial. This work addresses the selection of in the mirror descent (MD) sparse optimization framework such that a user-specified sparsity rate is achieved. Instead of treating this selection as a meta-optimization problem, which potentially requires expensive trial-and-error sweeps, our approach adapts online during training. We propose two schemes to perform the adaptation: one based on the magnitude of the largest subgradients, and the other based on the difference between the model's current and target sparsity. The first scheme can be interpreted as a dynamic sparse training cycle. A sparsity schedule further extends this control to intermediate targets, gradually sparsifying the model during training. We also provide a convergence analysis for adaptive in MD. Experiments on CIFAR-100, TinyImageNet, and ImageNet-1K show that both schemes reliably reach the final as well as the scheduled intermediate sparsity targets. Beyond controlling sparsity, the adaptive schemes match or exceed the accuracy of a tuned fixed at equal or higher sparsity in every evaluated setting, and clearly surpass it when combined with a sparsity schedule: with ResNet-50 on TinyImageNet, they reach 67.1% top-1 accuracy at 95% sparsity, compared with 64.6% for a fixed at only 90.4% sparsity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.