Diffusion Posterior Sampling for Fast Adaptation in Multi-task Nonlinear Contextual Bandits
Abstract
We study fast adaptation in multi-task nonlinear contextual bandits, where different tasks share the same reward structure but are characterized by distinct model parameters drawn from a common unknown prior distribution. Given data from past tasks, the goal is to rapidly adapt to a new task and minimize regret using only limited online interactions. Thompson Sampling (TS) is a popular approach for solving contextual bandits, maintaining a posterior over the model parameter that is updated each round using a hand-specified conjugate prior (e.g., Gaussian) and the observed rewards. However, such priors cannot capture the rich cross-task structure in multi-task settings, leading to misspecified posteriors and suboptimal exploration. To enable fast adaptation, we train a diffusion model on data from past tasks to learn a flexible transferable prior distribution over task parameters. In a new bandit task, parameters are estimated via a conditional reverse-diffusion process, where each step combines: (i) an unconditional drift from the diffusion prior, (ii) a likelihood-driven drift from the interaction history, and (iii) a noise term enabling randomized exploration. We instantiate this framework in two ways. DLTS integrates history into the diffusion prior at every reverse step to form a conditional posterior, from which approximate samples are drawn. DPSG applies a single history-guided gradient correction at each diffusion step. Theoretically, we formalize oracle TS and its diffusion counterpart, and show they are equivalent when the diffusion prior matches the true prior. Empirically, our methods are competitive with specialized baselines in linear settings and provide strong fast-adaptation performance in challenging nonlinear bandit environments, requiring substantially fewer online interactions to reach a target performance level.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.