Layer-propagation Equivalence Compression for Gradient-Based Data Selection in LLM Fine-Tuning
Abstract
Gradient-based data selection identifies useful training examples by comparing their gradients with those of a target validation set, which is crucial to large language model (LLM) supervised fine-tuning (SFT). Computing these gradients with an LLM scales poorly, making it impractical for multi-billion-parameter models, motivating the construction of smaller proxy models through compression for gradient-influence equivalence to preserve the gradient information used for data selection. However, existing compression objectives evaluate each weight matrix separately, without accounting for how the resulting errors propagate across layers; we find that these errors can persist and become amplified, potentially altering the gradient alignment scores used for data selection. To address this issue, we introduce Layer-propagation Equivalence Compression (LPEC), a method that accounts for error propagation when constructing a low-rank compressed proxy. Specifically, we perform propagation-aware direction ranking by considering both singular values and how removal errors propagate to other layers. We then perform joint rank allocation across weight matrices under a shared parameter budget, giving more rank to matrices where retaining additional components is expected to reduce compression error more effectively. Experimental results across different LLM families and evaluation tasks demonstrate that LPEC improves downstream SFT performance over existing influence-preserving compression methods. Anonymous code is available at https://anonymous.4open.science/r/LEPC-119E/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.