Confounded Linear Contextual Bandits
Abstract
This work studies **confounded contextual bandits**, where the reward is modeled as a linear function of the action, additively confounded by a **stochastic** term that is independent of the action. When the confounder is strong, it can dominate and attenuate the treatment effect induced by the chosen action, making reward inference from bandit feedback challenging. To address this, we draw on Neyman orthogonalization and propose a Thompson sampling based algorithm, INF-TS, which achieves regret over a horizon of , where denotes the dimension of the context-action tensor product and is the number of actions. We also prove a lower bound of for the confounded linear contextual bandit problem, showing that INF-TS attains the optimal regret rate in . Empirical results demonstrate that INF-TS outperforms existing methods designed for adversarial confounders in linear contextual bandits under both stochastic and adversarial confounder settings, particularly when the confounding signal dominates the treatment effect.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.