acceptodds
Under review as a conference paper at ICLR 2027

When LoRA Rank Matters in RLVR: Saturation-Gated Scaling and Off-Principal Initialization

Abstract

Does reinforcement learning with verifiable rewards (RLVR) need high-rank adapters? Under one fixed on-policy harness, we identify a saturation-gated rank-scaling law: on Qwen2.5-Math bases with RL headroom, accuracy rises with rank on the large benchmarks (7B MATH-500 39.2 → 72.6; 674-problem OlympiadBench 8.8 → 40.1 from base to rank-256), while the shared five-rank sweep on DeepSeek-R1-Distill-Qwen-1.5B is flat. The trained RLVR update has high effective rank (mean ≈ 0.8× nominal over 196 layers), consistent with a capacity bottleneck. Cyclic merge-reinitialization, orthogonal resets, and subspace consolidation fail to recover the high-rank benefit. Instead, OffPrin places capacity in the bottom right-singular subspace of each frozen weight through a function-preserving LoRA initialization. Its rank-32 adapters match standard rank-128 adapters with 4× fewer trainable parameters on both Qwen2.5-Math backbones (7B MATH-500 72.7 vs 72.2; OlympiadBench 41.1 vs 40.5); a principal-init control supports the subspace-specific effect. With full-scale, multi-seed training, OffPrin reaches 85.1 MATH-500 / 43.7 AIME-2024 on Qwen2.5-Math-7B under our internal harness. Under the standardized Sober harness it also surpasses strong full-parameter RLVR baselines, with a smaller adapter and optimizer-state footprint at comparable reasoning accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.