Hq-Recycle: Efficient Bilevel Optimization through Adjoint Recycling
Abstract
We introduce Hq-Recycle, a quasi-Newton method for bilevel optimization that reduces repeated work in implicit hypergradient estimation by recycling adjoint estimates across outer iterations. The method tracks the lower-level solution using limited-memory quasi-Newton updates and retains the previous adjoint estimate. A fresh directional finite-difference curvature probe produces a one-dimensional Galerkin initialization, followed by preconditioned corrections only when the estimated residual exceeds a prescribed tolerance. This procedure limits refinement to a small probe budget while accounting for changes in the current adjoint system. For exact symmetric positive-definite systems, we establish an energy-minimization property of the recycled initialization and an error bound governed by the previous residual and inter-iteration system drift. We evaluate the four-probe configuration on three label-corrupted linear classification tasks and grouped ridge regression against qNBO, SHINE, PZOBO, BOME, HOAG, and AID-CG, using validation-based tuning and five fresh evaluation seeds. In classification ablations, recycling reduces the mean number of adjoint probes from four to approximately 1.07 per outer iteration. On grouped ridge regression, Hq-Recycle achieves a 2.44-fold speedup over tuned qNBO in mean time to a prespecified validation target and lowers mean test loss by 20.2% under a 0.5-second solver budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.