acceptodds
Under review as a conference paper at ICLR 2027

Temper: Reading the Second Moment from the Riemannian Gradient in LoRA

Abstract

Low-rank adaptation trains a weight update as the product of two factors, and one update has many factorizations. Recent LoRA ptimizers are therefore built to be transformation-invariant: the step should not depend on which one is stored. One line secures that by stepping along the Riemannian gradient, the projection of the Euclidean gradient onto the manifold's tangent space and a function of the update alone. However, invariance fixes the direction of the step, not how far it moves along each coordinate. That distance is what a second moment sets, and the methods that step along the Riemannian gradient either carry none or read one from the Euclidean gradient, almost all of which lies where no rank- step can follow. To this end, we propose Temper, which accumulates one second moment per row and one per column of the Riemannian gradient itself, and never forms the dense matrix that reading it there is known to cost: a tangent vector is low rank, so both marginals follow exactly from two small Gram matrices.Temper is provably invariant and costs no more per step than an unpreconditioned one. On Llama-3-8B and Qwen-2.5-7B it leads the strongest baseline by to points on the commonsense average, GSM8K and all three HumanEval pass rates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.