Where to Reuse, What to Propagate: Transition Reconstruction for Cross-Domain Offline Reinforcement Learning
Abstract
Cross-domain offline reinforcement learning leverages abundant source data to alleviate target-domain data scarcity, but dynamics mismatch often induces severe negative transfer. While existing methods attempt to mitigate this by selectively reusing target-compatible source transitions, they inevitably force a harsh trade-off between preserving valuable state-action coverage and avoiding dynamics-induced bias. We analytically reveal that this bottleneck stems from treating transitions as indivisible units, which strictly couples where experience is reused with what original outcome is propagated. To break this inherent coupling, we propose shifting the paradigm from transition selection to transition reconstruction, thereby enlarging the admissible transition family. Specifically, our framework, T-RAFT, isolates target-supported source conditions via unbalanced optimal transport (UOT), and amortizes these cross-domain correspondences into a continuous outcome translator using conditional Flow Matching (FM). This decoupled mechanism allows the agent to safely inherit broad source coverage while seamlessly replacing biased next states with target-compatible dynamics. Empirical evaluations demonstrate that our approach consistently overcomes the coverage-bias dilemma, outperforming state-of-the-art baselines on standard benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.