acceptodds
Under review as a conference paper at ICLR 2027

SupportDICE: Replay-Supported Occupancy Optimization for Offline-to-Online Reinforcement Learning

Abstract

Offline-to-online reinforcement learning aims to reuse an offline dataset while adapting as new interaction data arrives. We study this regime from the DICE (distribution-correction estimation) occupancy-optimization perspective. The central difficulty is that the replay buffer keeps changing: it tells us which state-action pairs have been observed and provides samples for training, but raw replay counts need not themselves form a valid discounted occupancy for policy optimization. We propose SupportDICE, which separates these roles by constructing an augmented replay-supported MDP whose exits from the currently covered region are redirected to an absorbing state. The resulting update optimizes valid occupancies on the current replay support, while a full-support anchor lets newly observed, flow-feasible state-action pairs enter future updates. In finite MDPs, the exact population updates have zero duality gap under finite-divergence feasibility; the fixed-support offline phase recovers DICE improvement in the augmented MDP, while the support-expanding online phase satisfies a per-step comparator guarantee on the current replay-supported problem. We also extend the population analysis to dominated standard Borel spaces. A tabular square-grid example shows that support expansion can expose a better feasible occupancy, while replay-centered DICE remains biased toward the old replay distribution. On D4RL benchmarks, SupportDICE achieves the highest average final scores on Adroit and Kitchen among the evaluated methods, while remaining competitive on AntMaze.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.