What Sets a Block-Sparse Training Speedup? SM Count, the Dense Baseline and the Kernel, on Six GPUs
Abstract
Dynamic sparse training (DST) rarely delivers wall-clock speedups, because standard implementations multiply dense weights by masks and still execute dense products. We ask what sets the speedup once the sparsity is actually executed. Our instrument is BlockDST, which keeps a constant number of live blocks in every block-row so that our own load-balanced blocked-ELL kernels run all three per-layer products sparse, including a block sampled dense–dense matrix multiplication (block-SDDMM) for the weight gradient. On six NVIDIA graphics processing units (GPUs), the same kernel and shape run - faster than a dense General Matrix Multiply (GEMM) at 95% sparsity, and most of this spread comes from how fast the dense baseline runs on each card. Our kernel's throughput scales with streaming multiprocessor (SM) count clock, so a one-parameter model predicts it on a held-out card to within - at the reference shape (- across shapes). That scaling is not a property of block sparsity in general: cuSPARSE's Blocked-ELL kernel, timed on five of the same cards, does not follow it. It runs faster on an A100 than on an A40, while ours runs as fast, yet its speedup over dense still varies about from card to card. End to end the picture is narrower: on an L40, BlockDST reaches a fixed accuracy faster than dense (not significant with three seeds) with less memory, but it is slower than dense on an A40, slower than a dense model with the same number of live parameters, and, on a laptop GPU, - slower than the same block layout without rewiring. A sparse speedup should therefore be reported together with the kernel, the card's SM count and clock, and the dense throughput it was measured against.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.