FocusDreamer: Task-Directed Model-Based Reinforcement Learning Guided by Return Relevance
Abstract
Model-based reinforcement learning improves interaction efficiency through learned world models and imagination, but model bias can distort return differences among local behaviors and misguide policy optimization. When local alternatives are weakly supported by experience, their predictions may be biased toward familiar latent dynamics, causing real return advantages to shrink or even reverse in imagination. Therefore, under limited interaction budgets, it is more sensible to prioritize states where local behavioral distinctions matter more for return. We propose FocusDreamer, which learns state-level return relevance from completed real trajectories and uses it to focus exploration, model learning, and imagination without altering the original Dreamer objectives. Return relevance is characterized by criticality, which measures its relative strength, and polarity, which reflects the direction of return-predictive progression. Criticality determines where learning resources should be concentrated, while polarity modulates policy-side intervention. Across visual DMControl and Meta-World, FocusDreamer improves performance and sample efficiency, achieving a 53.1% average return improvement over DreamerV3 on DMControl.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.