acceptodds
Under review as a conference paper at ICLR 2027

Closing the Pareto-Response Loop: Preference-Aware Optimistic Selection for Multi-Objective Bilevel Learning

Abstract

Multi-objective bilevel learning (MOBL) coordinates conflicting upper-level objectives through lower-level responses. Yet prevailing methods typically use a response-first organization: they form the lower-level response first and only afterward reconcile its induced objectives or hypergradients according to user preferences, without directly feeding the Pareto mixture into the lower objective. This response-formation channel remains open-loop, although under non-convex or finite-step optimization, the reached response is highly path-dependent and fundamentally shapes the ensuing Pareto geometry. We introduce Preference-Aware Optimistic Selection (PAOS), which uses a selector-controlled lower objective for optimistic branch selection among non-unique lower solutions and preference-guided response shaping under finite optimization budgets. Specifically, a stateful response-selection variable integrates the preference-weighted Pareto mixture produced by the current subproblem and guides the subsequent lower-level response, thereby closing the Pareto-response loop. On an analytically tractable double-well problem, PAOS systematically selects the preference-compatible branch from near-boundary starts. Across data hyper-cleaning, few-shot bilevel learning, physics-informed neural networks, and medical image segmentation, PAOS achieves the strongest overall performance among the evaluated MOBL methods. These results establish preference-aware optimistic response selection as an actionable degree of freedom in multi-objective bilevel decision-making. Our code is available at https://anonymous.4open.science/r/PAOS-forreview/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.