Same Adapter Function, Different GRPO Dynamics: Factor-Scale Confounding in LoRA Optimization
Abstract
Low-Rank Adaptation (LoRA) factorizes each weight update into two trainable matrices, so the same adapter function admits multiple factor scales. Because standard optimizers update these factors separately, functionally identical initializations can receive different updates, confounding the attribution of gains from non-zero LoRA initializations. We study this ambiguity using function-preserving scalar gauges, an identical frozen-rollout control, and reciprocal factor learning rates. Under online Group Relative Policy Optimization (GRPO), shared-rate AdamW amplifies early adapter-product displacement by up to and separates trajectories from identical initial policies, while reciprocal rates substantially reduce this sensitivity across Qwen3 models up to 8B. A teacher-forced SFT control reproduces the local coordinate effect without rollout, reward, or advantage estimation. These results show that factor-scale sensitivity is an optimizer-level confound that is not unique to GRPO, while online policy feedback can make it consequential for early RL trajectories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.