acceptodds
Under review as a conference paper at ICLR 2027

Cross-Domain Energy-Guided Diffusion Generation for Off-Dynamics Reinforcement Learning

Abstract

Off-dynamics offline reinforcement learning seeks to learn a target-domain policy from a large source dataset and a limited target dataset under mismatched transition dynamics. Existing approaches such as reward augmentation and data filtering are constrained to the existing source dataset and thus have limited coverage beyond the collected dataset. Recent model-based methods address this issue by learning target-aware dynamics, but are limited to transition-level generation, which can lead to accumulated errors over long horizons. These limitations motivate trajectory-level generation for off-dynamics offline RL. We propose CEDGE, a Cross-domain Energy-guided Diffusion GEneration framework. CEDGE trains a trajectory diffusion model on source-domain trajectories and adapts the generated samples to the target domain through energy guidance. This guidance is derived from the mismatch between source and target trajectory distributions and decomposed into return, domain, and policy energies for high return, target-dynamics compatibility, and source behavior correction respectively. The resulting energy-guided trajectories are useful both for direct planning and as synthetic data for policy learning. Compared with previous methods, CEDGE performs target adaptation through energy guidance, and the source trained diffusion model can be reused across target dynamics without retraining. Our theoretical analysis studies generation error and downstream policy performance. Empirical results demonstrate that trajectory-level energy-guided generation improves diffusion planning under dynamics shifts and produces synthetic data that improves downstream target-domain policy learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.