acceptodds
Under review as a conference paper at ICLR 2027

RANK IS NOT A TRUST REGION: OPTIMIZER-INDUCED GEOMETRY IN LORA REINFORCEMENT LEARNIN

Abstract

Low-rank adaptation (LoRA) is widely used for reinforcement learning with verifiable rewards (RLVR), and a low rank is commonly expected to keep the policy close to where it started. We show that rank is not a trust region: in GRPO training of Qwen3.5-4B, rank-4 and rank-64 adapters that update the weights equally far differ by in on-policy KL, and by under an explicit KL penalty. Two updates matched in rank, norm, and row and column subspaces differ by in held-out KL. The reason is that policy KL depends on the direction of an update and not only on its size. Under standard initialization, the down-projection of LoRA barely moves, so its rows act as nearly fixed random probes of the gradient. As the gradient concentrates in directions to which the policy is most sensitive, a higher rank supplies more probes that capture this sensitivity, so the KL cost of a matched displacement grows with rank. Changing the probes tests this mechanism: probes aligned with the gradient reverse the rank ordering of curvature and, when KL stays local, reverse the rank ordering of KL as well. Directional Fisher sensitivity, which measures curvature along the realized update per unit norm, predicts KL on held-out sequences in two model families. Policy movement also does not determine whether an update helps, because KL is locally even in the update direction while utility change is odd. Stability in LoRA-based RLVR should therefore be monitored and constrained in distribution space, with rank chosen for capacity rather than as a trust-region parameter.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.