To Explore or to Exploit: Conservative Optimism with Pessimistic Baseline for Offline-to-Online Learning
Abstract
Offline-to-online learning begins with data collected before online interaction and then continues learning from newly acquired online feedback, raising a fundamental question: when should the learner trust the decision supported by offline data, and when is it worth paying the cost of exploration to improve upon it? The answer depends critically on the online horizon : committing to a pessimistic lower-confidence-bound (LCB) decision avoids costly early exploration but incurs large regret for long horizon if the decision is suboptimal, whereas optimistic upper-confidence-bound (UCB) exploration enables long-run improvement at the price of potentially large regret in the early stage. Yet choosing between these strategies in advance is generally impossible, since the relevant crossover depends on unknown instance gaps and the horizon may itself be unknown. We show that a simple cumulative-budget rule suffices to obtain direct baseline protection under the standard pseudo-regret criterion while preserving continued online learning. The resulting algorithm, Conservative Optimism with Pessimistic Baseline (COPB), maintains a cumulative performance certificate relative to the offline LCB decision: it follows the hybrid UCB recommendation when the certificate permits exploration, and otherwise plays the offline LCB baseline. For stochastic bandits, COPB simultaneously guarantees regret close to that of committing to the offline LCB decision, up to a prescribed tolerance, and hybrid UCB-type gap-dependent and gap-independent regret bounds with explicitly characterized conservative overhead. The bounds further capture how offline coverage reduces both online exploration cost and the cost of baseline protection. COPB is horizon-adaptive: its exploration-versus-exploitation behavior requires no horizon-dependent tuning and automatically shifts toward online improvement as evidence accumulates. We further extend the framework to episodic tabular reinforcement learning and validate it on synthetic and real-data experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.