Budget-Aware Layer Pruning for Large Language Models
Abstract
Layer pruning is attractive for large language model deployment because it yields practical gains in latency and throughput. The key challenge is selecting which layers to remove under a pruning budget , where denotes the number of layers to remove. Many existing methods use a single, budget-independent ranking and select the top- layers, imposing a nested-prefix constraint across pruning budgets. Through exhaustive search on LLaMA3.1-8B, we show that optimal pruning sets can be non-nested across budgets, exposing a limitation of fixed global rankings. Motivated by this finding, we shift layer scoring from intrinsic importance to a budget-aware continuation value and formulate layer pruning as a budget-conditioned finite-horizon Markov decision process. We approximate this value using Double Dueling DQN, enabling a single scorer to construct pruning sets conditioned on both the target budget and the current pruning context. Experiments demonstrate better language modeling quality and downstream accuracy than strong depth-pruning baselines at the same pruning budget, and greater robustness under aggressive pruning. Code is available at https://anonymous.4open.science/r/Budget-Aware-Layer-Pruning-B2CD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.