Learned Warmup and Target-Conditioned Handoff for Contextual Bandits
Abstract
Many contextual bandit deployments are streams of short-lived instances: per title, campaign, participant, or item pool. Each instance starts with little local history, so standard online algorithms pay the exploration cost repeatedly, even when aggregate traffic is large. We study prior-data fitted networks (PFNs) as a way to amortize this cold-start phase across related bandit tasks, deploying a model through standard forward passes conditioned on an instance's observed history. We instantiate this idea as Bandha (Bandit Amortization with Neural Decision Handoff and Adaptation), using three mechanisms. First, the sampled PFN environment supplies full-feedback targets for all arms during training, while the model operates only on the censored history available online. Second, a permutation-equivariant per-arm state lets the learned update rule apply across different numbers of arms. Third, an outer Thompson-sampling controller selects the action-generating policy (AGP) used to produce informative training histories. Together, these mechanisms train the model from teacher-augmented synthetic episodes while its deployed state updates only from censored histories, rather than carrying deployment state between instances. We evaluate the construction on fixed-dimensional contextual bandits. Empirically, the learned update rule transfers across both arm counts and datasets. A single model trained on , where is the number of arms, is evaluated at K\in {3,5,7,10,13,15,20\} using cumulative pseudo-regret. At horizon 200, the model reduces this regret by 4.2% to 33.6% relative to linear upper confidence bound (LinUCB), and improves over linear Thompson sampling (LinTS). At horizon 1000, it reduces regret by 21.6% to 54% relative to LinUCB in paired evaluations on the same instances at . On 17 real-world classification datasets, the same model remains useful zero-shot after fixed preprocessing. When another model, such as LinUCB, must take over, a warmup prefix optimized for its downstream continuation (target-conditioned handoff) improves on an untargeted prefix, while unconditioned or misaligned prefixes show where amortized exploration can hurt downstream learners.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.