acceptodds
Under review as a conference paper at ICLR 2027

Layer-wise Sensitivity-aware Sparsity Allocation for Efficient LLM Inference

Abstract

Large Language Model (LLM) inference presents substantial computational challenges when executed on commodity hardware, thereby necessitating the development of efficient acceleration techniques. While existing approaches predominantly focus on uniform compression strategies, they neglect the heterogeneous sensitivity patterns exhibited across different transformer layers. In this paper, we introduce the Adaptive Sparsity Allocation Framework (ASAF), which integrates rotation-based low-bit quantization with layer-wise adaptive sparsity allocation. The framework comprises two sequential phases with a dynamic programming strategy. In Phase 1, coarse-grained optimization determines the optimal number of layer groups and narrows sparsity rate search intervals. In Phase 2, fine-grained optimization determines precise consecutive layer allocation and exact sparsity rates within each group. The joint optimization of layer grouping decisions and sparsity rate assignments creates a combinatorial explosion in the solution space, rendering brute-force approaches computationally prohibitive; we employ a dynamic programming strategy that decomposes the exponential search space into manageable subproblems across both phases, achieving practical computational efficiency while guaranteeing optimality within the discretized two-phase search space. Experiments on the Llama-2 family, with additional memory and zero-shot results on Llama-3, reveal that our proposed framework sustains average benchmark accuracy degradation within 1%, concurrently achieving up to 3.63 prefill acceleration over an FP16 baseline and 12.63% memory reduction over the quantized baseline, on a single NVIDIA RTX 3090 GPU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.