Implicit Thompson Sampling via Imagined Rewards
Abstract
Thompson sampling (TS) is an efficient algorithm for addressing the exploration-exploitation trade-off in sequential decision-making. However, its application to modern auto-regressive predictive models (e.g., LLMs and TabPFN) is limited, since it requires sampling from the posterior of an explicit Bayesian model. To address this challenge, we propose Imagined Reward Thompson Sampling (IRTS) as a family of TS-style algorithms, based on the insight that in-context learning with autoregressive models provides predictive inference for an implicitly defined Bayesian model. IRTS uses autoregressive models to sample imagined rewards autoregressively per action, and acts greedily w.r.t. either their averaged rewards (IRTS-A), or the predictive mean reward conditioned on the observed and imagined rewards (IRTS-C). We provide a regret analysis and prove that IRTS-A and IRTS-C provide a “sandwich” of exploration behaviour over TS, where both algorithms approach TS when . With finite , IRTS-A over-explores due to over-estimation of posterior uncertainty, while IRTS-C under-explores by interpolating between greedy and TS. Experiments across multi-armed and contextual bandits, with analytic, TabPFN and LLM-based in-context learners, demonstrate that IRTS achieves efficient TS-style exploration with implicitly-defined Bayesian models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.