Revealing Local Nonlinearity Amplification for Large Language Model Compression
Abstract
The growing computational cost of large language models (LLMs) has driven increasing interest in structured compression to reduce inference overhead. Existing structured compression methods primarily assess layer importance or redundancy, but rarely characterize how errors introduced by compression propagate through subsequent computation. We observe a consistent layer-wise pattern in LLMs: across inference activation distributions, layers exhibit distinct levels of local nonlinear deviation, with many being well approximated by affine mappings. Meanwhile, the response of downstream computation to layer outputs exhibits stable distributional patterns across layer positions. We therefore introduce Local Nonlinearity Amplification (LNA), a quantitative measure that combines a layer's local nonlinear deviation with the downstream response to its output change, providing a direct estimation of the degradation risk associated with layer compression. Guided by LNA-based layer selection, we propose ReLoNA, which adaptively determines an affine-operator replacement strategy based on each layer's local nonlinearity. After layer replacement, ReLoNA applies Operator-Aligned Recovery to reduce the approximation errors introduced by replacement and maintain consistency with the original computation. Across 18 benchmarks and multiple models, ReLoNA outperforms existing structured compression methods at 25% compression ratios, retaining 94% of dense-model performance on multiple-choice benchmarks and achieving better performance on generative benchmarks, with gains of up to 2 over existing methods. Moreover, LNA-guided pruning consistently yields substantially lower degradation than existing criteria under the same compression ratio.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.