acceptodds
Under review as a conference paper at ICLR 2027

Implicit Thompson Sampling via Imagined Rewards

Abstract

Thompson sampling (TS) is an efficient algorithm for addressing the exploration-exploitation trade-off in sequential decision-making. However, its application to modern auto-regressive predictive models (e.g., LLMs and TabPFN) is limited, since it requires sampling from the posterior of an explicit Bayesian model. To address this challenge, we propose Imagined Reward Thompson Sampling (IRTS) as a family of TS-style algorithms, based on the insight that in-context learning with autoregressive models provides predictive inference for an implicitly defined Bayesian model. IRTS uses autoregressive models to sample imagined rewards autoregressively per action, and acts greedily w.r.t. either their averaged rewards (IRTS-A), or the predictive mean reward conditioned on the observed and imagined rewards (IRTS-C). We provide a regret analysis and prove that IRTS-A and IRTS-C provide a “sandwich” of exploration behaviour over TS, where both algorithms approach TS when . With finite , IRTS-A over-explores due to over-estimation of posterior uncertainty, while IRTS-C under-explores by interpolating between greedy and TS. Experiments across multi-armed and contextual bandits, with analytic, TabPFN and LLM-based in-context learners, demonstrate that IRTS achieves efficient TS-style exploration with implicitly-defined Bayesian models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.