RCA: Fisher-Residual Low-Rank Adaptation for RLVR Reasoning
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training paradigm for improving language-model reasoning. Existing RLVR work has primarily studied reward sources, rollout sampling, KL regularization, etc. In this paper, we study a complementary question: how does RLVR act on the local parameter space of an adapted model, and which adapter update directions provide higher reward improvement per unit policy-change cost? We approach this question through the local Fisher geometry induced by KL-aware RLVR. Our results show that, under a reward–curvature separation condition, updates in the bilateral Fisher-residual channel space can achieve higher reward-per-curvature efficiency than equal-rank selectors that include high-curvature modes. Inspired by this conclusion, we introduce RCA (Residual-Channel Low-Rank Adaptation), a Fisher-guided Low Rank Adapter (LoRA) method for RLVR reasoning. RCA first estimates layerwise Fisher input and output channels with a calibration-time Kronecker-Factored Approximate Curvature (KFAC) approximation. The resulting projectors define a residual-channel low-rank update constraint. Empirically, RCA improves reasoning reliability across pass@1–pass@128 compared with full-finetuning RL and representative LoRA variants, while maintaining a LoRA-sized trainable parameter footprint.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.