Dynamic Gradient Modeling with Quality-Guided Masking for Adaptive LLM Fine-Tuning
Abstract
Gradient masking reduces redundant parameter updates during large language model (LLM) fine-tuning, but fixed global retention policies overlook variation across layers and training stages. We introduce DGMM, a dynamic gradient modeling algorithm with quality-guided masking that adaptively allocates gradient retention budgets across layers. DGMM integrates three complementary signals—directional sign bias, temporal stability, and inter-layer correlation—into a layer-wise gradient quality score. These scores are smoothed with an exponential moving average and used to allocate retention budgets within stage-aware bounds, while gradient magnitudes determine which elements are retained within each layer. Operating exclusively on optimizer-side gradients, DGMM preserves the model architecture, forward computation, training objective, and underlying optimizer. Experiments on task-specific and general-domain benchmarks show improvements of up to 30.8% over the strongest baseline on individual benchmarks. Controlled sweeps evaluate masking rates of up to 90% and reveal no change in profiler-reported training FLOPs. These findings support gradient-quality-guided budget allocation as an effective mechanism for adaptive gradient masking and underscore the distinction between sparse parameter updates and reduced training computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.