Factorization Depth as a Regularization Dimension for LoRA
Abstract
Existing deep low-rank formulations largely treat depth as a fixed design choice; we investigate whether factorization depth can instead serve as a regularization dimension for Low-Rank Adaptation (LoRA), beyond trainable parameter count and structural-rank capacity. To characterize this role, we analyze a mergeable -factor parameterization, with controlling factorization depth, through its representation cost, induced optimization geometry, and idealized singular-mode dynamics. These analyses motivate a finite-horizon mechanism hypothesis: deeper factorization may selectively delay weak singular modes, promoting spectral concentration at moderate depths but hindering task fitting when taken too far. We test this hypothesis through four complementary approaches: capacity controls, depth ablations, spectral diagnostics, and synthetic experiments. Across these experiments, the resulting accuracy–rank–loss patterns support the hypothesized trade-off between spectral concentration and task fit, rather than a monotonic benefit from increasing depth. Finally, to probe the mechanism behind this effect, we show that a nuclear-norm intervention recovers only part of the performance gain observed at moderate depths, suggesting that the depth effect is not fully explained by explicit nuclear-norm regularization in the tested settings. Overall, our results support viewing factorization depth as a promising, non-monotonic regularization dimension for LoRA, while motivating further study of whether this phenomenon generalizes across optimizers, backbone models, and tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.