acceptodds
Under review as a conference paper at ICLR 2027

Predictive Signature Diversity for Model-Based Offline Reinforcement Learning

Abstract

Offline reinforcement learning (offline RL) aims to learn policies from static datasets, but suffers from distributional shift when evaluating out-of-distribution actions. Model-based approaches mitigate this issue using learned dynamics models, yet remain highly sensitive to model uncertainty, especially in low-coverage regions. Existing approaches fail to adequately capture such uncertainty, particularly in terms of worst-case discrepancies between models. We propose a predictive, policy-independent framework for model selection based on predictive signature diversity. Each model is represented by its multi-step predictions on shared anchor inputs, enabling comparison of behaviors in a common space. We periodically select a compact subset of accurate models by maximizing worst-case separation under this representation, yielding a dynamically updated ensemble for rollout generation and policy learning. Our approach decouples diversity estimation from policy rollouts and directly captures model disagreement in uncertain regions. We provide empirical results demonstrating improved robustness over a broad range of baselines, including model-free, model-based, ensemble-based, and mutual-information-based methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.