Trust-Region On-Policy Sparse Optimization for Pruning Reasoning Language Models
Abstract
Reasoning language models incur high inference costs because they generate long chains of thought before producing an answer, making model compression techniques such as pruning increasingly important. However, current pruning approaches typically cause severe performance degradation at high sparsity, leaving effective reasoning model pruning an open problem. In this work, we identify a key challenge in pruning reasoning models: pruning alters the prefixes—the prompts and partially generated sequences—encountered during generation, causing the fixed, off-policy calibration data to become increasingly mismatched with the evolving sparse model. To tackle this challenge, we introduce SCOUT, a novel sparse optimization framework that alleviates the distributional mismatch by leveraging on-policy data from the evolving sparse model, effectively curbing reasoning performance drop. At its core, SCOUT dynamically refines the sparsity mask through iterative pruning and reviving, while constraining each update within a KL divergence-based trust region to bound policy drift. We further establish a stationarity bound for an idealized form of this alternating procedure, where the trust-region budget and rollout staleness govern the resulting convergence neighborhood. Empirically, we showed that SCOUT substantially outperforms existing pruning methods at high sparsity across different Qwen3 models in various scales.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.