Decentralized Zeroth-Order Gradient Tracking For Bilevel Optimization
Abstract
In many learning and decision-making problems, the quality of a decision depends on the outcome of an underlying optimization process. For example, selecting a model's regularization strength requires training the model under that choice and assessing its performance on validation data. Model training and hyperparameter selection therefore form two nested optimization problems. When organizations jointly train a shared model while keeping data local and operating without a central coordinator, both problems must be solved through neighbor communication, motivating decentralized bilevel optimization. Such coordination becomes more difficult when participating systems expose only loss values or have insufficient memory for differentiation. These constraints motivate a zeroth-order approach that uses function evaluations at both levels. The challenge extends beyond estimating local gradients: agents must recover how the jointly trained model responds to hyperparameter changes from noisy local evaluations, despite incomplete training and disagreement across the network. To recover this response without derivatives, we propose Decentralized Zeroth-Order Bilevel Optimization (DZBO), which optimizes both levels through stochastic function evaluations and communication between neighbors. For a possibly nonconvex outer objective and uniformly strongly convex local lower-level objectives, DZBO achieves an rate for the average expected squared gradient norm at network-averaged iterates after outer updates. Attaining accuracy requires function-value queries per agent and communication rounds, including all inner and auxiliary work, for fixed problem and network parameters. Experiments illustrate the feasibility of our proposed algorithm under heterogeneous data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.