Understanding, Stabilizing and Accelerating High-Ratio N:M Sparse Training in Deep Reinforcement Learning
Abstract
Row-wise sparsity has emerged as a promising approach to accelerate matrix-matrix multiplications. However, the behavior of this type of sparsity remains under-explored in deep reinforcement learning (RL). Although magnitude-based sparsity implementations perform well at lower sparsity levels, they fail to maintain performance as sparsity increases. In this paper, we systematically investigate, stabilize, and accelerate high-ratio row-wise : sparse training in continuous-control RL. Under static 1:8 sparsity, randomly sampled initial mask topology has little predictive power for final return, whereas lower Fraction of Active Units (FAU) in the dense critic generally accompanies smaller sparse-dense return gaps. Under dynamic sparsity, we identify a three-stage actor saturation collapse and a mask-update shock in the critics, and counter them with scale-tuned initialization and transient replay-ratio increases. With increased width and a higher replay ratio, stabilized 1:8 agents achieve higher mean returns than same-width dense agents in 13 of 18 environment-buffer configurations. We further introduce time-multiplexed : training, which co-trains adjacent discrete patterns through shared parameters while keeping each subnetwork deployable at inference. On Walker2d, both subnetworks of a 3.5:8 schedule, trained with 12.5% fewer active weights than fixed 4:8, exceed independently trained 3:8 and 4:8 networks in mean return. Finally, our custom CPU sparse kernels provide 2.17-3.95 speedups for the forward pass of 1:8 layers and, combined with an approximate active-weight-only Adam optimizer, 1.34-1.83 end-to-end training speedups over an MKL-backed dense-kernel reference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.