Generative Support Realignment for Cross-Domain Offline Reinforcement Learning
Abstract
Cross-domain offline reinforcement learning learns a target policy from pre-collected source and target datasets with different dynamics. When target data are scarce, effectively compensating for their limited coverage using source data remains challenging due to the discrepancy between domains. We propose *Target-aligned Coverage Expansion (TCE)*, which leverages source states to generatively realign and expand the limited target support while controlling the generation error induced by this expansion. We further derive a performance gap bound that characterizes the interplay between generation error and source–target dynamics gap, providing theoretical guidance for effective coverage expansion and source utilization. Across diverse cross-domain environments, TCE consistently outperforms state-of-the-art baselines, while further analysis empirically supports the guidance provided by our theoretical findings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.