Understanding Sparse Fine-Tuning: Lessons and Insights from a Unified Empirical Study
Abstract
Sparse fine-tuning updates a small subset of task-critical neurons and has been reported to outperform full fine-tuning (FFT) while resisting catastrophic forgetting. Existing methods, however, differ in the selection criterion, module, schedule, and neuron-update ratio, and why sparse fine-tuning works remains poorly understood. We select eight representative sparse fine-tuning methods and conduct a unified empirical study on mathematical reasoning and code generation, using FFT as a reference, establishing a performance leaderboard of sparse fine-tuning methods. We further ablate the key design choices underlying sparse fine-tuning to provide a comprehensive analysis of this paradigm. We characterise the role of the sparse masked update: it acts as a capacity allocator that concentrates learning on task-critical neurons, and as a regulariser that pins the remaining parameters to their pre-trained values. This dual role mitigates the trade-off between task performance and anti-forgetting in FFT: sparse methods improve over FFT by points on average while retaining – of the base model's general capability across all tested families, scales, and tasks. We further find that: (1) The advantage of sparse fine-tuning over FFT grows with the capability of the base model; (2) Task-relevant neurons reside in both the MHA and the FFN of the backbone, so the selection location has only a marginal effect; (3) The optimal neuron-update ratio is method- and task-dependent, making adaptive neuron selection essential; (4) Dynamic mask regeneration trades general capability for a modest task gain, at a recurring compute and memory cost. Our analysis clarifies the design choices of sparse fine-tuning and identifies future research directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.