Beyond Non-Zero Initialization: Investigating Update Balance in Structured LoRA Initialization
Abstract
Low-rank adaptation (LoRA) fine-tunes pretrained models through two small factor matrices. Structured initializers choose a nonzero adapter using spectral or contextual information, but even the same initialized adapter can be represented by factor pairs that learn differently. We identify their relative scale as an overlooked degree of freedom: reciprocal rescaling leaves the initial adapter and frozen residual unchanged while redistributing the factors’ contributions to an update. Our analysis shows how the selected structure determines this allocation and distinguishes correctable imbalance from limits that scalar rescaling cannot overcome. For orthogonal spectral factors, we derive an exact population scaling rule under isotropic Gaussian inputs and single-sample SignSGD. This leads to spectral and activation-based scale choices; InitCal complements them with reversible AdamW probes that measure the optimizer’s response. Across Qwen3 and Llama models, calibration improves the four-task average in all nine structured model–initializer settings, with gains in 35 of 36 task-level comparisons. Paired controls show that update magnitude alone does not explain these improvements. Factor-scale allocation thus provides a practical way to improve structured initialization without changing the adapter it selects. The code is available [here](https://anonymous.4open.science/r/initcal-7406).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.