POP: Prior-Fitted First-Order Optimization Policies
Abstract
Adaptive first-order optimizers rely on fixed, hand-designed mechanisms for incorporating optimization history, limiting how flexibly they can adapt to a new objective. We introduce POP, a prior-fitted first-order optimizer whose learned action is deliberately restricted to trajectory-conditioned coordinate-wise step sizes along the descent direction. POP is meta-trained via reinforcement learning on millions of synthetic optimization problems sampled from a task-independent compositional prior with Gaussian process components, using scale-invariant policy inputs to support transfer across objectives. Despite training only on low-dimensional synthetic problems, POP generalizes to longer horizons, higher-dimensional objectives, and unseen problem distributions. On a benchmark of 43 diverse optimization functions, POP significantly outperforms per-function tuned gradient-based baselines and transfers zero-shot to practical optimization tasks on real-world datasets. Our work provides a proof of concept that transferable optimization behavior can be learned through trajectory-conditioned step-size adaptation alone, and that such behavior can emerge from meta-training exclusively on objectives sampled from a synthetic generative prior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.