Activation Sparsity as Training–Inference Co-Design
Abstract
Activation sparsity can accelerate transformer inference when a specialized kernel can skip computation associated with zero activations. We ask how broadly and aggressively train-time sparsification should be applied to obtain such exploitable sparsity without excessive quality loss. Using matched Pythia-family pretraining runs, we add fixed-threshold nonlinearities at selected sites in the network forward pass, systematically vary activation pressure, and connect the resulting sparsity to a specialized-kernel execution. We find that extending sparsification across more activation sites can substantially increase model-wide sparsity without reducing latency. At 14M parameters, more than doubling model-wide sparsity from 12.71% to 27.48% slightly increases latency. Conditional kernel analysis identifies the FFN hidden activation \(h\) and attention-output projection input \(z\) as the most useful sparsification sites. Pressure applied only at \(h\) also increases sparsity at \(z\), despite \(z\) receiving no direct pressure. Broader pressure produces stronger optimization conflict at less responsive sites. These observations motivate a targeted intervention that achieves lower validation loss than the dense control while reducing latency from 0.652 to 0.573 ms. Scaling analysis confirms the susceptibility to sparsification of but reveals an increasing overhead from the specialized kernel. Our results suggest that sparsification is best treated as a training–inference co-design problem: rather than maximizing sparsity globally, training should induce specific exploitable structure that matches the inference system.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.