Beyond Instantaneous Gradients: Momentum-Guided Adaptive Gradient Masking for Sparse Fine-Tuning of LLMs
Abstract
Sparse fine-tuning of large language models (LLMs) requires selecting the parameters that contribute most to downstream adaptation under a limited update budget. Existing gradient-based selection methods rely primarily on gradient statistics from the current optimization stage, making their importance estimates susceptible to mini-batch noise, outliers, and short-term fluctuations in optimization. We propose Momentum-Guided Adaptive Gradient Masking (MAGM) to address this issue. MAGM uses an exponential moving average to integrate gradient information across training steps, yielding a history-aware importance estimate that captures sustained parameter contributions. These importance scores guide layer-wise budget allocation and within-layer parameter selection under a fixed global update budget. Experiments show that MAGM achieves the highest average scores among the compared methods in code generation, mathematical reasoning, and general-domain evaluation. On DeepSeek-Coder-Base-6.7B and MISTRAL-7B, MAGM improves the average code-generation score by 1.63% and 0.52%, respectively, over the best-performing algorithm. In mathematical reasoning, MAGM achieves average performance comparable to the best-performing baseline on both MISTRAL-7B and LLAMA3-8B. On general-domain tasks, the gains over the best-performing algorithm are 2.0% for supervised fine-tuning and 2.1% for direct preference optimization. These results indicate that MAGM enables more stable parameter importance estimation and more effective budget allocation under a limited parameter update budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.