SPECTRA: Spectral Stable Learning Rate Prediction for Low-rank Adaptation
Abstract
We propose **SPECTRA**, a gradient-spectral framework for predicting stable learning rates for LoRA fine-tuning. Although LoRA substantially reduces the cost of full-parameter SFT, its performance remains highly sensitive to learning-rate selection. Existing studies offer valuable scaling laws, heuristics, and empirical insights, but they do not directly predict an appropriate learning rate for a specific task and LoRA configuration, leaving practitioners reliant on costly hyperparameter sweeps. **Methodologically**, SPECTRA uses a calibrated log-linear predictor to map the initialization-time gradient spectral decay and LoRA rank to a stable, near-optimal learning rate. **Theoretically**, we derive the local stability threshold for rank-constrained optimization, identify the most fragile module as the shared-rate bottleneck, and establish Adam’s gradient-scale invariance in the idealized limit. Under a local monotonic relationship between spectral decay and curvature, faster spectral decay reduces the stability threshold. **Experimentally**, SPECTRA achieves an **88.5%** success rate, outperforming all sweep-free baselines by over **9.7** points and recovering **89%** of the coverage of an exhaustive per-dataset grid search at **4%** of its time cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.