acceptodds
Under review as a conference paper at ICLR 2027

Does Energy-Aware Sparsity Allocation Help Dynamic Sparse Training? Leverage, Execution Back Ends, and Stable Regrowth

Abstract

Dynamic sparse training (DST) is motivated by lower training cost, and a natural refinement is to allocate sparsity by energy, pruning hardest where computation is most expensive. We test whether this helps, measuring GPU energy with hardware counters in randomized blocks on three execution back ends. An elementary identity bounds what any layer-wise allocation can save at a fixed number of weights: the saving is the drop in the average cost per weight when weights move to cheaper layers. This leverage is zero for masked dense kernels, as used by most DST code, and for the linear layers of Transformers, and it is large in CNNs, whose early layers cost 100-200x more per weight than their late layers. A score based on each layer's total energy captures it only on AlexNet; allocating by energy per weight captures about half of it. The measurements show why this leverage still does not pay off. Under cuSPARSE kernels, allocation effects are real and predicted within 2.8 points by per-layer costs, but energy is not monotone in density: a uniform network with twice the weights uses 44% less energy than one at the common density of 0.1 and is more accurate. Under channel-structured execution, sparse training finally costs less than dense training (0.40x on VGG-16), but at equal measured energy neither energy-aware allocation nor its reverse beats simply lowering the density of a uniform allocation. Energy-aware regrowth fares worse: it follows a replicator dynamic that collapses the network into one layer (-9.8 points on AlexNet) unless it is anchored, which we prove and confirm fixes it. On the hardware we measured, the execution back end and the density set the energy of sparse training, not the layer-wise allocation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.