acceptodds
Under review as a conference paper at ICLR 2027

Tangent Stein Bandits under Predictable Common Shifts

Abstract

Common shifts in candidate features leave action rankings unchanged but can bias raw moment estimates of the reward direction. We study this problem in single-index contextual bandits with an unknown monotone link and Gaussian candidates, allowing the common shift to depend on past observations. Within-set centering removes the shift from action comparisons but correlates the candidates, invalidating the independent-winner score. We show that the selected centered action retains an exact Gaussian moment in the tangent space of the current direction. This identity enables learning without evaluating the link or its derivative for candidate sets containing at least two actions, including candidate counts where the full centered score has a boundary or second-moment obstruction. With finite conditional noise variance, certified scale bounds, and sufficient initialization, a coordinate-clipped tangent epoch algorithm achieves polylogarithmic-in-horizon high-probability pseudo-regret under fixed positive derivative-information and schedule constants. The sharper guarantee follows from a quadratic conditional-mean regret relation specific to Gaussian candidates. In 18 held-out simulation settings, continual updates improve mean pseudo-regret over an independently tuned, valid centered explore-then-commit learner; shared initialization comparisons isolate the benefit of subsequent updates. Six long-horizon comparisons also favor tangent updates over a tuned persistent uniform-exploration policy. Matched perturbations distinguish the distributional conditions for the tangent identity from the derivative information required for effective learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.