Prox: Training-Free FFN Activation Sparsification in LLMs via Approximate Intermediate-State Salience
Abstract
Feed-forward networks (FFNs) dominate memory traffic and computation in large language model (LLM) inference, making them a primary target for activation sparsification. However, existing training-free methods suffer substantial model-quality degradation at high FFN sparsity due to limitations in their channel-selection strategies. We observe that the SwiGLU intermediate state provides a highly effective channel-selection signal, but obtaining it requires costly dense computation. To address this, we present Prox, a training-free framework for sparse SwiGLU FFNs based on the insight that sparse execution needs only the channel mask induced by the intermediate state, which can be constructed from the magnitude ranking of its entries (i.e., channel salience) rather than their precise values. Prox works in two stages: Stage1 approximates the ranking via input sparsity and quantized proxy weights to construct the channel mask; Stage2 computes the selected channels exactly using full-precision weights, enabling sparse execution of all three projections. Extensive evaluations across ten LLMs from six model families show that Prox consistently outperforms training-free baselines at various sparsity levels, achieving up to a end-to-end decoding speedup at 70% FFN sparsity, and is compatible with quantization and sparse attention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.