acceptodds
Under review as a conference paper at ICLR 2027

Exploring What Matters: Value-Aware Exploration for Dreamer

Abstract

Dreamer improves data efficiency by learning a world model from real experience and training its actor and critic on imagined trajectories. These imagined updates depend on the experience collected, making efficient exploration essential under a limited interaction budget. Predictive disagreement can waste interaction on dynamics with little effect on return. Immature value estimates can also overlook useful experience, limiting task-relevant model learning and policy improvement in imagination. We propose Value-aware Exploration (VEX), which directs exploration in DreamerV3 through value-priced predictive uncertainty. VEX evaluates alternative transition predictions through a shared task readout of reward, continuation, and future value, measuring their disagreement in terms of task return. We derive a Bayesian allocation priority under imperfectly observed task sensitivity, recovering uncertainty-only exploration in the uninformative limit. VEX’s exploration score directly weights imagined actor updates to guide experience collection. Across three continuous-control tasks, both VEX operating points improve aggregate sample efficiency and final return over DreamerV3 and same-backbone exploration baselines. Ablations demonstrate the contributions of its key components.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.