A New Approach to Bilevel Optimization: A Value-Centric Perspective via a Distributional Follower
Abstract
Classical bilevel optimization represents the lower-level response by a selected minimizer . This optimizer-centered formulation can be unstable in overparameterized or multimodal settings. Moreover, when multiple lower-level minimizers exist, it requires a response-selection rule that may be difficult to justify in practice. We propose an optimal-value-centered formulation that replaces the selected minimizer by a Gibbs distribution over lower-level solutions. Equivalently, this distribution solves an entropy-regularized lower-level problem over probability measures, while the upper level minimizes the expected outer objective. This representation motivates Value-Centric Hypergradient Descent (VCHD) algorithm, a double-loop method that samples from the decision-dependent Gibbs follower using mini-batch SGLD and updates the outer variable with a sample-based hypergradient estimator. We establish nonasymptotic convergence without assuming lower-level convexity and show that VCHD reaches an -stationary point with sample complexity. We conducted a thorough study of the empirical performance of the VCHD algorithm: a quadratic example examines how temperature and inner-loop length affect its performance, while a double-well example shows that VCHD can approach the correct outer solution even when the lower-level landscape is nonconvex and contains local minima. Results on the hyper-data-cleaning application further suggest that our optimal-value-centered hypergradient can preserve informative update signals when classical bilevel methods suffer from overfitting and vanishing hypergradients.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.