acceptodds
Under review as a conference paper at ICLR 2027

LoRA Initialization Geometry and Learning-Rate Transfer

Abstract

LoRA parameterizes weight updates as a product of low-rank matrices, making the geometric constraints imposed by this structure important for understanding its training dynamics. Initialization determines the tangent space from which learning begins and the local curvature, constraining the directions of weight changes and their departure from the initial tangent space. We analyze how this geometry, which varies with the initialization, interacts with the learning rate to shape training. First, we use the initial spectrum and cumulative movement to analyze subspace rotation and local departure from the initial tangent space. Our analysis indicates that geometric differences relative to the initialization alone do not readily explain the performance attained. We further investigate whether that weight update magnitudes provide a more direct explanation for the performance attained. From this analysis, we derive a learning-rate scaling rule that matches the magnitudes of first-order weight updates across initialization scales. We establish a criterion for assessing learning-rate ranges over which both training loss trajectories and final evaluation performance are comparable across settings. Using this criterion to guide learning-rate selection, we show that existing LoRA initialization methods can achieve similar evaluation performance when their learning rates are tuned suitably. We also identify the potential for learning-rate transfer across different initialization methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.