acceptodds
Under review as a conference paper at ICLR 2027

Partially Linear Contextual Bandits

Abstract

In contextual bandit problems, traditional methods for modeling the context–reward relationship typically employ either fully parametric or nonparametric approaches. Parametric approaches are computationally efficient and simple to implement, but they rely on restrictive assumptions that may not hold in complex real-world scenarios. Nonparametric methods offer greater flexibility, but may not fully leverage clear underlying patterns that could otherwise improve prediction. In this work, we bridge these two paradigms by proposing a partially linear model for the expected reward, in which a designated subset of contextual features influences the reward linearly while the remaining features act on it through an unknown function belonging to a reproducing kernel Hilbert space. We introduce PartLinUCB, an algorithm designed for this partially linear reward model. PartLinUCB naturally encompasses both linear and kernelized bandits as special cases, yielding a unified regret guarantee that recovers the established rates for each when the context–reward relationship is fully linear or fully nonparametric, respectively. We further establish that PartLinUCB is nearly minimax optimal for the radial basis function (RBF) kernel by deriving a lower bound that matches our regret upper bound up to logarithmic factors. Empirical studies demonstrate that PartLinUCB outperforms representative baseline methods and highlight the practical utility of the partially linear approach.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.