acceptodds
Under review as a conference paper at ICLR 2027

Same Adapter Function, Different GRPO Dynamics: Factor-Scale Confounding in LoRA Optimization

Abstract

Low-Rank Adaptation (LoRA) factorizes each weight update into two trainable matrices, so the same adapter function admits multiple factor scales. Because standard optimizers update these factors separately, functionally identical initializations can receive different updates, confounding the attribution of gains from non-zero LoRA initializations. We study this ambiguity using function-preserving scalar gauges, an identical frozen-rollout control, and reciprocal factor learning rates. Under online Group Relative Policy Optimization (GRPO), shared-rate AdamW amplifies early adapter-product displacement by up to and separates trajectories from identical initial policies, while reciprocal rates substantially reduce this sensitivity across Qwen3 models up to 8B. A teacher-forced SFT control reproduces the local coordinate effect without rollout, reward, or advantage estimation. These results show that factor-scale sensitivity is an optimizer-level confound that is not unique to GRPO, while online policy feedback can make it consequential for early RL trajectories.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.