Layer Pruning as a Non-Cooperative Game: Conditional Exchanges for Large Language Models
Abstract
Structured layer pruning reduces the computational and memory costs of large language models, but the effect of removing a layer depends on which other layers remain. Static rankings may overlook this dependence, while surrogate-based approaches introduce an additional prediction stage. We formulate fixed-budget layer selection as a finite non-cooperative game with a shared utility and develop a conditional exchange procedure. Starting from standalone layer scores, the method evaluates exchanges that restore one removed layer and remove one retained layer. At each iteration, it accepts the exchange that most reduces the Kullback-Leibler divergence between dense and pruned output distributions on calibration data. The search terminates when no permitted exchange improves the objective beyond a numerical tolerance, without weight updates, surrogate training, or Monte Carlo Shapley estimation. Experiments across multiple model families, pruning budgets, and language-modeling benchmarks show lower perplexity than the comparison method in most evaluated configurations, while zero-shot results reveal budget-dependent trade-offs. Matched inference measurements show comparable latency, throughput, and allocated GPU memory. In a separate wall-clock evaluation on a seven-billion-parameter model, our method is about faster overall than the comparison method across four pruning budgets. These results support conditional exchange as a practical approach to selecting layers while accounting for their interactions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.