acceptodds
Under review as a conference paper at ICLR 2027

Post-Hoc Paired-Coordinate Reduction in LoRA-RL Reasoning Adapters

Abstract

Parameter-efficient reinforcement learning makes reasoning post-training practical, yet completed low-rank adapters are typically retained at the rank chosen before or during optimization. Existing adaptive-rank methods mainly modify training itself, leaving open whether a completed LoRA-RL factorization contains removable capacity. We study this question through post-hoc paired-coordinate selection: checkpoint-derived scores rank shared LoRA factor coordinates, a fixed Top-k set is sliced consistently across modules, and the original scaling is preserved without retraining or using downstream task outcomes for selection. On a rank-32 LoRA-RL adapter, evaluations across MATH500, AIME25, and AMC23 at reveal substantial but task- and budget-dependent retention relative to random selection and per-module SVD reconstruction. Validity diagnostics further show that successful endpoints do not imply intrinsic or representation-invariant coordinate importance. As a cross-family test, applying the Energy selector to a separately trained DoRA-r32 adapter retains MATH500 and AMC23 performance at k = 16 while AIME25 decreases, showing that post-hoc reduction transfers beyond the original LoRA parameterization but remains task dependent. These results identify usable post-hoc redundancy while bounding its interpretation to the stored factorization and evaluation protocol.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.