PALORA: PARTIAL LOW-RANK ADAPTATION FOR LANGUAGE MODEL POST-TRAINING
Abstract
Low-rank adaptation (LoRA) reduces trainable parameters in language model post-training but applies the same update parameterization across target matrices. We introduce Partial Low-Rank Adaptation (PALORA), which directly trains selected Transformer projection matrices and retains LoRA elsewhere. Its probe–select–reinitialize procedure begins with short LoRA-only runs of Group Relative Policy Optimization (GRPO). Complementary binary masks yield absolute adapter-gate gradients of the GRPO loss on fixed held-out rollout groups. For each matrix, we normalize mean sensitivity by its standard deviation across probe questions plus a numerical stabilizer, then average scores across probe seeds. The lowest-scoring matrices receive direct updates: an empirical allocation rule rather than a diagnosis of insufficient adapter rank. After discarding probe weights, final GRPO training restarts from the original checkpoint with fresh adapters and a fixed allocation. We evaluate instruction-tuned Qwen2.5 models at 1.5B, 3B, and 7B parameters on five mathematical reasoning benchmarks. The 1.5B configuration directly trains 14 of 196 candidate matrices, retaining rank-16 LoRA at the remaining 182. At validation-selected checkpoints, PALORA improves mean accuracy across the five benchmarks over the reported LoRA baselines by 2.15, 3.18, and 4.26 percentage points, respectively. These single-seed comparisons use nonidentical parameter budgets and do not isolate placement effects.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.