acceptodds
Under review as a conference paper at ICLR 2027

Post-Selection-Valid Conformal Guarantees over Continuous Preferences for Multi-Objective Reinforcement Learning

Abstract

Many multi-objective reinforcement learning applications operate in a decision-support mode: a system first reveals a set or curve of trade-offs, and a user then chooses one operating point for execution. Existing guarantees are usually pointwise and protect a preference fixed in advance; they need not remain valid when the preference is selected after the full curve has been revealed. We study this release-then-select problem and propose CERTPREF, which publishes a preference-indexed certified lower-bound curve for the realized normalized scalarized return of the policy that is actually executed. CERTPREF calibrates one-sided conformal bounds at finitely many frozen anchor policies, makes the anchor guarantees simultaneous, extends them to all preferences using a trajectory-level Lipschitz envelope, and executes the same anchor policy that defines the displayed bound. This yields a distribution-free marginal post-selection guarantee for a single deployment query over a continuous preference space. Experiments across multi-objective control benchmarks show that CERTPREF controls post-selection violations below the target level, often conservatively, while retaining informative lower bounds. Pointwise baselines can be tighter, but they are not designed to control selection after the curve is revealed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.