PRESTO: PeRspective-based Entropy-Sparsity Training Optimization
Abstract
Pruning and quantization make a neural network smaller, and when the goal is storage size, a lossless compressor is eventually applied (Han et al., 2016) to a model that is usually not trained to be compressible. We introduce PRESTO, a quantization-aware training method that adds to the loss a proxy for the compressed size of the weights, leaving the activations in floating point. This proxy is the value of a convex inner problem over the quantization grid learned by LSQ (Esser et al., 2020), written in perspective form, i.e. the tightest convex relaxation of the decision to keep a weight. A Lagrangian relaxation makes that problem separable, and an envelope argument gives a (sub)gradient of its value. To our knowledge, PRESTO is the first method to jointly balance sparsity and the codebook symbol distribution arising from quantization within a single convex inner problem, rather than through separate terms that act independently, as in HEMP+LOBSTER (Tartaglione et al., 2021) and LilNetX (Girish et al., 2023). On ImageNet, with sizes measured on the whole checkpoint after lossless coding, PRESTO reaches on ResNet-18, on ResNet-50, on ViT-B/16, on DeiT-Small and on EfficientNet-B0, above full-precision accuracy on the two ResNets and less than accuracy points below it on the other three; on five of the six networks, we hold an operating point that no published method we compare against betters on both axes. On AlexNet, prune-first methods remain ahead, a deficit we attribute to sparsity allocation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.