Context-Aware Preference Bandits with Heterogeneous Human Feedback
Abstract
In context-aware preference learning, an online learner infers user preferences from sequential feedback. We study this problem under the contextual bandit framework, where each arm is associated with an unknown feature vector. However, context dictates not only the optimal arm but also alters the evaluation scale itself, in that user feedback across scenarios is inherently uncalibrated and highly heterogeneous. To capture these effects, we model the linear utility of each arm via a context-dependent affine calibration, parameterized by a multiplicative scale and an additive offset. We propose a phased elimination algorithm that combines pilot-estimation sample splitting with a one-step Newton correction. For pointwise rating feedback, we group contexts using score binning and learn each group separately, achieving a near-optimal regret, where parameterizes the tail decay rate of the score bounds. For pairwise comparison feedback, the additive offset cancels out exactly, yielding a near-optimal regret for any . We further establish lower bounds for both feedback protocols, matching the common estimation term in our upper bounds. Experiments on synthetic data and benchmark datasets validate our theoretical findings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.