acceptodds
Under review as a conference paper at ICLR 2027

Task Residuals Are Low-Rank: Efficient Pruning and Recovery of Large Language Models

Abstract

Structured pruning reduces the inference cost of large language models by removing complete architectural units. However, existing methods either rely on task-agnostic importance estimates that overlook task-specific structural requirements or repeatedly evaluate all candidate structural blocks for each downstream task, resulting in substantial adaptation overhead. Through a systematic block-wise masking study, we empirically identify a previously overlooked structure in block importance: it can be decomposed into a generic component shared across tasks and task-specific residuals concentrated in a compact low-rank subspace. Building on this finding, we propose ReTaPER, a residual subspace-guided framework for task-aware pruning with efficient recovery. ReTaPER first learns a low-rank residual basis from block-masking responses. For a new task, it probes only a small set of informative anchor blocks, reconstructs the complete task residual within the learned subspace, and combines it with generic importance to derive a task-specific pruning order. After pruning, the reconstructed residual further guides the allocation of a fixed LoRA rank budget, assigning greater adaptation capacity to retained blocks that are more sensitive to the target task. Extensive experiments across nine tasks demonstrate that ReTaPER achieves state-of-the-art pruning performance both with and without recovery. Code is available at https://anonymous.4open.science/r/ReTaPER-ICLR2027-4DD3/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.