Learnable Adaptive Spectral Reparameterization for Improving LLM Pretraining
Abstract
Recent advances suggest that flatter and better-conditioned weight spectra are closely associated with improved LLM pretraining and generalization, motivating increasing interest in shaping the spectral geometry of model parameters. However, existing approaches typically rely on prescribed spectral transformations, whose predefined mappings impose strong priors on spectral evolution and may restrict the flexibility required during training. Moreover, different layers and parameter types can favor distinct spectral geometries, making a uniform transformation potentially suboptimal across the model. To address this limitation, we propose Learnable Adaptive Spectral Reparameterization (LASR), which enables individual weight matrices to learn their preferred spectral transformations from the training objective. Specifically, LASR constructs a learnable family of spectral mappings through a truncated K-th-order expansion of the reciprocal square-root function, with trainable coefficients controlling the contributions of different polynomial bases. This parameterization spans diverse feasible spectral amplification behaviors while ensuring bounded spectral reshaping. By jointly optimizing these coefficients with the model, each weight matrix adaptively learns its spectral transformation according to its training dynamics. Consequently, LASR turns spectral reshaping from a prescribed, model-wide operation into a weight-specific learnable mechanism. Extensive experiments across pretraining, downstream evaluations, and spectral analyses consistently demonstrate the effectiveness of LASR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.