Value-Directed Exploration for Efficient Model-Based Reinforcement Learning
Abstract
Efficient exploration remains a central challenge in reinforcement learning, particularly when environment interactions are costly. Curiosity-motivated algorithms guide exploration using model uncertainty as an intrinsic reward, however resolving errors in rarely visited or return-insensitive regions does not necessarily improve policy performance. Consequently, such strategies may allocate exploration effort inefficiently with respect to the task objective. To address this, we propose a value-directed optimistic exploration strategy that evaluates model uncertainty through its effect on future value, thereby prioritizing information acquisition in regions where reducing uncertainty can most affect which actions or policies are preferred. Our analysis yields a value-weighted information gain that can be substantially smaller than standard information gain when value sensitivity concentrates on a low-dimensional subset of the dynamics, giving regret bounds whose leading term tightens with this reduction and otherwise recovers value-agnostic scaling. Empirically, value-directed exploration improves performance in environments with task-irrelevant distractors while matching strong baselines on standard benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.