acceptodds
Under review as a conference paper at ICLR 2027

Understanding Update Selection in Low-Rank Fine-Tuning

Abstract

Low-rank adaptation (LoRA) is commonly evaluated through final benchmark scores, which do not isolate the effects of individual training updates. We study whether additional evidence improves one-step update selection and what explains the resulting gains. Holding the model and optimizer history fixed, we compare candidate updates and measure their one-step cross-entropy reduction on held-out examples. Locally improving updates can harm held-out data, and proposals with matched effective-weight displacement norms can have different held-out gains. A frozen selection rule using additional task-balanced examples achieves mean gain 0.002071 on previously unused examples with Llama-3.2-1B, compared with -0.001090 for full updates and -0.000011 for current-loss selection. A Qwen3-4B-Base replication also shows higher mean gain with additional evidence. Reassigning choices within trajectories reduces gain while preserving candidate frequencies and acceptance rates, showing that overall preferences and fewer updates do not fully explain the benefit. Preserving acceptance locations retains much of the gain. On Qwen, simpler evidence-based policies and controls that preserve training step also retain much of the benefit. These results provide a controlled account of the gains from one-step update selection in low-rank fine-tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.