acceptodds
Under review as a conference paper at ICLR 2027

When More Objectives Make Identification Easier: The Role of Preferences and Feedback

Abstract

Best-arm identification (BAI) aims to identify the optimal arm using as few samples as possible. While most existing work focuses on a single objective, many real-world applications involve multiple criteria, which motivates the study of BAI in multi-objective bandits. Intuitively, multi-objective identification seems harder, since the learner must identify optimal arms in a higher-dimensional reward space; however, the additional objectives also provide richer feedback that can reduce the required exploration. In this paper, we characterize when and why multiple objectives can make BAI easier. For Gaussian bandits, we establish a characteristic-time lower bound for fixed-confidence identification that holds uniformly across preference orders. Through explicit comparisons under Pareto, linear, and lexicographic preferences, we reveal three phenomena: (i) identifying a larger optimal set can require fewer samples than identifying a unique optimum; (ii) the advantage of vector feedback over coordinate feedback depends fundamentally on the preference order; and (iii) multi-objective identification can require fewer samples than even the easiest corresponding single-objective problem. Finally, we propose Order-Aware Track-and-Stop (OATS), an algorithm that is correct at any prescribed confidence level and asymptotically matches the lower bound.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.