acceptodds
Under review as a conference paper at ICLR 2027

Transfer Learning with Neural UCB

Abstract

We study transfer learning in neural contextual bandits, where labeled data from a source task are available before interaction with a target task whose reward function may differ. We propose Transfer-NeuralUCB, which uses a predictor trained on the source task as the center for target adaptation and transfers source uncertainty through a design-weighted certificate for reward mismatch. The algorithm combines a mismatch-adjusted source UCB with an online UCB learned from target feedback and selects actions using their pointwise minimum. Given an RKHS transfer certificate and finite-width neural tangent kernel conditions, we establish simultaneous confidence guarantees along the source and target training iterates. The resulting regret bound is controlled by the smaller of the cumulative source and online confidence widths. Our analysis identifies residual tangent complexity as the regularization cost of adapting the source predictor and characterizes, through inverse-kernel alignment and scale, when the source prediction is simpler to correct than the target reward is to learn from scratch.Experiments on synthetic and classification-derived bandits show that Transfer-NeuralUCB improves on NeuralUCB with a cold or warm start and on LinUCB in most settings, that the gain comes from source-centered residual learning and persists over rounds, and that it turns into a loss under strong source–target mismatch.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.