Learning from Successes with Multi-Step Implicit Values
Abstract
Discovering a successful trajectory does not ensure that a reinforcement learning agent learns effectively from it. In long-horizon, sparse-reward tasks, evaluating the current policy can assign low values even to states on successful trajectories, weakening the learning signal from rare successes. Building on Implicit Q Learning, we propose multi-step implicit value estimation, which combines expectile-based value learning with recursive temporal propagation. Rather than evaluating the behavior policy alone, the formulation targets an implicit policy that favors higher-valued actions while retaining unbiased expectations over rewards and transitions. We establish contraction of a joint action–state value operator and show that its multi-step compositions share a unique fixed point equal to the expected-return value of this implicit policy under the original MDP. We introduce \IQ, a coupled backward recursion for action and state values, and \IV, a value-only specialization for deterministic environments who recovers generalized advantage estimation (GAE) at the symmetric expectile. Experiments show improvements over PPO with GAE on multiple sparse- and dense-reward tasks, with complementary gains from exploration bonuses. Ablations highlight the importance of multi-step propagation and the versality of our methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.