RIFT: Reward Improvement Field Transfer for Cross-Domain Offline RL
Abstract
Learning from limited target data is a central challenge in cross-domain offline reinforcement learning. Although source data can provide additional experience, differences in dynamics can make these transitions unsuitable for the target domain. We propose Reward Improvement Field Transfer (RIFT), a method that instead uses source data to learn how behavior changes from low reward to high reward. RIFT encodes source and target transitions in a shared latent space, pairs low- and high-reward source transitions using optimal transport, and fits an affine reward-improvement field to their displacements. We apply this transformation to target transitions and decode the results to augment the target dataset. Our key assumption is that related domains may differ in their dynamics while still sharing a similar reward-improvement structure. We bound synthesis error under shared improvement geometry, identify conditions favoring target anchoring over source-data transplant, and derive a conditional connection to downstream return. Experiments on locomotion tasks with changes in gravity, morphology, and friction show that RIFT improves aggregate performance over the existing baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.