acceptodds
Under review as a conference paper at ICLR 2027

Demystifying Cross-Domain Policy Adaptation with World Models

Abstract

Cross-domain policy adaptation seeks to reuse source-domain knowledge under dynamics shifts with limited target-domain interaction. Model-based methods address this problem by adapting source-pretrained world models, but their apparent sample-efficiency advantage may partly arise from value transfer: model-based methods often retain source-trained value functions, while model-free baselines learn critics from scratch. When we give model-free baselines the same source-trained critic, they initially adapt faster than their Dyna-style counterparts. Further analysis identifies that dynamics ensembles and prior-guided data collection jointly improve Dyna-style adaptation, allowing it to outperform the model-free baseline on average. Guided by these findings, we propose \DomainADEPT, which improves model-based planning both with a dynamics ensemble to improve its accuracy and with “tethered” proposal policy fine-tuning to limit policy drift during data collection. Across four contact-rich manipulation tasks and one locomotion task, \DomainADEPT achieves the highest average final performance among the evaluated methods under matched target-domain interaction budgets, improving average success by 25 percentage points over the strongest baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.