acceptodds
Under review as a conference paper at ICLR 2027

GramLoRA: Cross-Scale Optimization Trajectory Transfer for LoRA Fine-Tuning

Abstract

Parameter-efficient fine-tuning (PEFT), particularly low-rank adaptation (LoRA), has become a standard approach for adapting large language models without updating the full parameter space. Existing LoRA variants mainly improve the parameterization, initialization, or optimization of low-rank updates, while adaptation is typically learned independently for each target model. We ask a complementary question: can the optimization trajectory of a smaller model guide LoRA adaptation of a larger one? We introduce GramLoRA, a cross-scale framework that transfers the optimization trajectory of a smaller same-family source model to guide LoRA adaptation of a larger target model. We fully fine-tune the source model and convert intermediate parameter residuals into target-compatible rank-\(r\) factors. The final source state provides a task-informed, identity-preserving initialization, while intermediate states provide stage-aware directional guidance without changing the downstream objective. Across Qwen2.5-7B, Gemma4-E4B, and Ministral-8B on eight commonsense reasoning tasks, GramLoRA consistently improves adaptation under limited target-model training. With 20% of Commonsense170K for 0.5 epoch, GramLoRA outperforms HiRA by 2.2, 4.6, and 2.9 average points, respectively, and also generalizes to natural language understanding and mathematical reasoning. Its advantage diminishes with extended target-model training, suggesting that transferred source-model optimization information is particularly effective for early target-model adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.