acceptodds
Under review as a conference paper at ICLR 2027

Offline-to-Online Distributional Reinforcement Learning

Abstract

Distributional reinforcement learning (DRL) has demonstrated strong performance and improved training stability by modeling the full distribution of returns. Despite this progress, we observe that in the offline-to-online setting, DRL suffers severe performance degradation at the onset of online finetuning. We trace this instability to quantile penalization during offline training, which may violate the intrinsic monotonic ordering of quantile values and induce substantially biased value estimates. To address this issue, we propose a monotonic quantile positional embedding and a monotonic quantile value network, which provably preserve quantile ordering under offline penalization. By enforcing this structural property, our method mitigates quantile estimation bias and stabilizes policy transfer from offline pretraining to online fine-tuning. We further introduce a quantile-guided exploration strategy that leverages the learned return distribution as an uncertainty signal to facilitate efficient online exploration. Experiments on the D4RL benchmark demonstrate that our method substantially reduces performance degradation at the onset of online interaction and consistently achieves superior final performance compared with existing offline-to-online reinforcement learning methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.