acceptodds
Under review as a conference paper at ICLR 2027

Value-Directed Exploration for Efficient Model-Based Reinforcement Learning

Abstract

Efficient exploration remains a central challenge in reinforcement learning, particularly when environment interactions are costly. Curiosity-motivated algorithms guide exploration using model uncertainty as an intrinsic reward, however resolving errors in rarely visited or return-insensitive regions does not necessarily improve policy performance. Consequently, such strategies may allocate exploration effort inefficiently with respect to the task objective. To address this, we propose a value-directed optimistic exploration strategy that evaluates model uncertainty through its effect on future value, thereby prioritizing information acquisition in regions where reducing uncertainty can most affect which actions or policies are preferred. Our analysis yields a value-weighted information gain that can be substantially smaller than standard information gain when value sensitivity concentrates on a low-dimensional subset of the dynamics, giving regret bounds whose leading term tightens with this reduction and otherwise recovers value-agnostic scaling. Empirically, value-directed exploration improves performance in environments with task-irrelevant distractors while matching strong baselines on standard benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.