Beyond Scalar Effective Learning Rates: Update Geometry under Parameter Symmetries
Abstract
Parameter symmetries preserve a model's function but can change the speed and direction of its updates. We describe these changes through *effective learning geometry*, using a *quotient update map* that recovers angular effective learning rates (ELRs) for scale-invariant models as a special case. For strictly rectangular value–output (VO) factors of full rank, we prove that Euclidean factor responses admit one common scalar on the entire product tangent space exactly for orthogonal basis changes. Probes with zero optimizer history in a transformer with 123.6M parameters reveal distinct compatibility structures: AdamW responds to both basis orientation and nonuniform scaling, whereas Muon is nearly compatible with orthogonal changes but remains sensitive to nonorthogonal scaling. Directional mismatch dominates the residual, with little contribution from variation in optimal rates across batches. Starting from identical functions with corresponding non-VO updates, fitting a scalar directly to the full output response reduces AdamW's residual from 0.192 to 0.102 and Muon's from 0.082 to 0.064. For both optimizers, the fitted residuals remain larger than those under exact update transport. Optimizer history attenuates the VO mismatch for AdamW and has a small, mixed effect for Muon. Paired pretraining trajectories show how the tested scalar compensations alter loss and prediction dynamics. Symmetry, update direction, and optimizer history thus determine what a change in learning rate can absorb.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.