Initialization Shapes the Computational Regime of Adaptive Optimization
Abstract
Optimization is central to modern machine learning, and its efficiency directly shapes the cost of model training. However, initialization is usually viewed as affecting only the starting point and transient trajectory, rather than the asymptotic convergence rate. We show that this separation can fail for adaptive quasi-Newton optimization. Parameter initialization alone can select distinct asymptotic computational regimes with the learning objective, optimizer family, and line-search hyperparameters held fixed. On a multidimensional regularized pinball-loss family, we identify an infinite hierarchy of such regimes for the Broyden–Fletcher–Goldfarb–Shanno (BFGS) method and its limited-memory variant (L-BFGS), and develop a trace-to-complexity analysis that links recurring line-search decisions to asymptotic progress, evaluation cost, and stability. The analysis reveals an exact tradeoff between contraction and sensitivity to initialization, and we further show that each exact regime persists over any fixed finite horizon within an open full-dimensional neighborhood. Extensive experiments, including multidimensional and real-world learning settings, show that the predicted structure is observable in finite precision and persists beyond the exact analytic construction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.