Pilot Priors: Safely Gating Warm Starts for Contextual Bandits
Abstract
Warm-starting contextual bandits has become an increasingly attractive method for reducing early-stage regret. Warm start data can come from historical logs, transfer pipelines, or be generated synthetically via simulators or large language models. However, the quality of such priors is often unknown: priors aligned with target tasks can reduce downstream regret, while misaligned priors can increase regret relative to cold starts. This paper studies the pre-deployment setting in which a single candidate prior is available, but only a short amount of real target-task data can be collected before source trust is fixed for the next deployment phase. We propose the Pilot Priors framework, which uses a short target-task pilot to certify a candidate warm prior relative to a cold-start baseline. The framework produces accept, reject, or inconclusive decisions from certified bounds on prior error, and satisfies a correctness guarantee for hard decisions about prior-error ordering. We further derive a sufficient condition for decisiveness through a margin-versus-width condition that links pilot informativeness to prior quality. Finally, for the inconclusive regime, we introduce a soft trust-calibration rule that interpolates between warm and cold starts rather than forcing a brittle all-or-nothing decision. The framework is evaluated in both a synthetic setting and two deployment scenarios based on real-world data. We show that the certificate becomes decisive as pilot information increases in synthetic experiments and detects harmful LLM-generated priors in one real-world setting. In another real-world setting, the certificate remains appropriately conservative while the soft calibration reduces harm from a reward-permuted null.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.