acceptodds
Under review as a conference paper at ICLR 2027

G-TUNE: MEMORY-EFFICIENT FINE-TUNING VIA COMPONENT-LEVEL GRADIENT DITCHING

Abstract

Fine-tuning large language models is essential for adapting them to downstream tasks and user preferences, but it is prohibitively memory-hungry. Intermediate activations, account for the dominant share of this memory, making memory-efficient fine-tuning critical. We perform memory-efficient fine-tuning at the granularity of individual components, selecting for each input only the attention heads and MLP blocks that matter, rather than treating entire tokens uniformly. This requires overcoming two challenges: (1) scoring and selecting the right components for each input, for which we introduce a gradient-guided importance score computable in a single forward-backward pass, and (2) keeping training stable despite per-instance selection, which we address via a balanced mask assignment procedure that casts sparsity-level assignment as a multi-objective search using NSGA-II. G-TUNE, a memory-efficient fine-tuning method via component-level gradient ditching. Across three models, Llama3.2-3B, Llama2-7B, and Qwen2.5-7B, G-TUNE surpasses the state of the art in both accuracy and memory efficiency, and remains effective when combined with LoRA and QLoRA.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.