acceptodds
Under review as a conference paper at ICLR 2027

Context-Aware Preference Bandits with Heterogeneous Human Feedback

Abstract

In context-aware preference learning, an online learner infers user preferences from sequential feedback. We study this problem under the contextual bandit framework, where each arm is associated with an unknown feature vector. However, context dictates not only the optimal arm but also alters the evaluation scale itself, in that user feedback across scenarios is inherently uncalibrated and highly heterogeneous. To capture these effects, we model the linear utility of each arm via a context-dependent affine calibration, parameterized by a multiplicative scale and an additive offset. We propose a phased elimination algorithm that combines pilot-estimation sample splitting with a one-step Newton correction. For pointwise rating feedback, we group contexts using score binning and learn each group separately, achieving a near-optimal regret, where parameterizes the tail decay rate of the score bounds. For pairwise comparison feedback, the additive offset cancels out exactly, yielding a near-optimal regret for any . We further establish lower bounds for both feedback protocols, matching the common estimation term in our upper bounds. Experiments on synthetic data and benchmark datasets validate our theoretical findings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.