Round-Trip UCB: Forward-Inverse Prediction for Neural Contextual Bandits
Abstract
Effective exploration in neural contextual bandits requires informative uncertainty estimates with low computational overhead. We introduce Round-Trip UCB, which measures uncertainty through the round-trip residual of forward prediction and inverse reconstruction. A forward head estimates an outcome from a candidate's learned representation, and an inverse head reconstructs that representation from the predicted outcome. The round-trip residual supplies an exploration bonus before feedback is observed. This design reuses the representation and prediction already needed for reward estimation, adding only inverse reconstruction and a residual calculation. It requires no candidate-specific parameter gradients or parameter-space uncertainty updates during action selection. In a noiseless fixed neural tangent kernel model, we prove under RKHS and candidate-density conditions that a computable multiple of the jointly trained residual bounds the prediction error with high probability, yielding polylogarithmic expected regret. On MNIST and Statlog Shuttle, finite-network Round-Trip UCB has and lower mean cumulative regret than NeuralUCB and about to less wall time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.