Ghost Hypergradients: Value-Gap Oracles for Optimistic Bilevel Learning
Abstract
Optimistic bilevel learning is fragile when the lower-level problem has many minimizers: differentiating an arbitrary selected response can follow a solver-dependent branch rather than the validation-preferred one. We study a value-function alternative based on a vanishing validation tilt. Let be the lower-level training value and the value after adding times the validation loss to the lower objective. The scaled value gap converges, under standard compactness and continuity conditions, to the optimistic objective that minimizes validation loss over the training-optimal set. Building on value-function penalties and equilibrium propagation, we make this perturbation identity algorithmic: by differencing conservative adjoint selections for and , we obtain ghost hypergradients, exact conservative oracles for the ideal finite-tilt surrogate that avoid differentiating a set-valued argmin map. Read at one shared random outer point, their exact scaled difference is almost surely a surrogate gradient, and a certificate error bound lets solve residuals control the computed direction. We prove local surrogate-error bounds, sharpened to for , projected-stationarity convergence, and inexact-solve bounds, exposing the bias–accuracy tradeoff in . On an exact flat-valley diagnostic, Ghost preserves the optimistic direction while selected-response gradients are misaligned. In our empirical study, Ghost achieves the best final test accuracy/runtime tradeoff among the compared bilevel oracles. These results support value gaps as practical first-order oracles for nonsmooth optimistic bilevel learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.