PRISM: Principal-Residual Activation Sparsity via Structure-Aware Budget Allocation for Efficient LLM Inference
Abstract
Training-free activation sparsity accelerates Large Language Model (LLM) inference by skipping computations on selected activation channels, without requiring additional training. Existing methods usually construct sparse computation patterns by selecting channels in the original activation coordinates. However, under high sparsity, the few retained channels may fail to cover sufficient output-relevant information, leading to large projection errors. Beyond the issue of channel selection, existing sparsity allocation methods commonly estimate module sensitivity with a single metric, overlooking structural dependencies inside Transformer blocks and the impact of error propagation. To address these limitations, we propose PRISM, a training-free activation sparsity framework with principal-residual computation and structure-aware sparsity allocation. PRISM first preserves low-dimensional principal components of linear-layer inputs through a compact dense path, and then applies sparse computation only to the residual activation. This dual-path design reduces the information loss caused by direct channel selection. PRISM further allocates sparsity budgets according to structural dependencies in Attention and MLP modules, avoiding overly aggressive sparsification of sensitive components. Experiments on representative LLMs show that PRISM better preserves model performance under high sparsity while providing practical inference speedup. At 75% sparsity, PRISM improves average reasoning accuracy by 3.7 points over the strongest sparse baseline and achieves up to 1.82 end-to-end generation speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.