Adaptive optimization improves scaling in quadratic neural networks
Abstract
We study the scaling laws of adaptive training algorithms for quadratic neural networks, and characterize scaling exponents parameterized by target complexity and model size. We demonstrate that an adaptive algorithm utilizing accumulated prediction errors can significantly improve the scaling behavior compared to fixed-learning-rate methods. Notably, this approach achieves statistically optimal exponents at any time-step without requiring manual learning rate tuning. Finally, we empirically demonstrate that the acceleration phenomenon due to adaptive learning rate generalizes: similar scaling improvements are observed under general activation functions and when employing practical adaptive algorithms that accumulate squared gradients.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.