Prune Better Before Recovering Harder: Trajectory-Aware Pruning and Recovery for Reasoning Language Models
Abstract
Reasoning language models are expensive to serve because they generate long autoregressive reasoning traces. Structured layer pruning can reduce this cost by shortening model depth, but criteria based on hidden-state similarity may remove layers that are important for reasoning. We propose trajectory-aware pruning and recovery, a preservation-first framework for compressing large language models. For each candidate layer, we temporarily replace it with an identity mapping and measure how much the model's next-token predictions change on fixed reasoning trajectories generated by the original model. We prune layers with the smallest behavioral effect, then use on-policy distillation to correct the remaining distribution shift on contexts visited by the pruned model. This preserves reasoning behavior before compression and reserves distillation for the errors that remain. Experiments on mathematical reasoning, code generation, and open-ended generation show that our method consistently outperforms SOTA method demonstrating that better layer selection reduces the burden on recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.