acceptodds
Under review as a conference paper at ICLR 2027

COE: Mitigating Degradation while Improving Fine-tuning in Offline-to-Online RL

Abstract

In offline-to-online RL, agents often experience performance degradation in the early stages of fine-tuning. There is some evidence that offline (batch) RL algorithms can mitigate this degradation, but can also learn too slowly during fine-tuning. The majority of progress in offline-to-online RL has been on achieving faster fine-tuning, but with insufficiently controlled performance degradation. We propose a new algorithm, called Cautious Online Expansion (COE), that combines a conservative batch algorithm and a faster-learning online algorithm, using off-policy estimation to gradually hand off control. We show the resulting algorithm mitigates degradation while allowing for fast learning online, compared to several batch algorithms and offline-to-online algorithms. We provide theoretical support for the backstepping approach in the algorithm and how the policies are evaluated to effectively decide on the hand-off. We additionally ablate several design decisions in COE, providing empirical support for each component.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.