Rethinking LLM Pruning via Test-Time Scaling
Abstract
Model pruning aims to compress large language models while preserving their capabilities, yet its effects on reasoning remain incompletely understood. We revisit expert, FFN and layer pruning through test-time scaling on mathematical reasoning and coding benchmarks. Pass@ curves distinguish success frequency from finite-budget reasoning coverage. Pruned models can preserve or exceed the original model's coverage even when sampling efficiency declines. Expert pruning more consistently preserves accuracy and coverage than FFN or layer pruning in the core matched-rate comparisons. Expert selection further affects both coverage and the sampling and token costs of reaching it, with advantages that vary across backbones. As removal increases, retained coverage can require substantially greater generated-token work, while severe pruning can also degrade generation stability. A mean-accuracy-matched uniform-loss reference does not fully explain the observed coverage losses. Models pruned from larger backbones can also retain broader reasoning coverage than native smaller models, although these comparisons do not match total parameters or training compute. Together, these findings provide an empirical guide to pruning that considers reasoning coverage, sampling efficiency and generation stability jointly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.