acceptodds
Under review as a conference paper at ICLR 2027

Bridging Reward Regression and Projected Bellman Equation with Successor Features

Abstract

Successor features (SF) decompose the Q-function into a successor feature map and a task vector, enabling efficient transfer in continual and goal-conditioned reinforcement learning. The Q-function estimate depends critically on how the task vector is estimated. We formalize the two estimators used in the literature—reward regression and the projected Bellman equation—prove that they generally yield different solutions, and characterize when they coincide. We unify them through an interpolation parameter and establish a stability–approximation tradeoff: reward regression stabilizes off-policy updates, while moving toward the solution of projected Bellman equation tightens the worst-case approximation guarantee. Building on this, we develop SF-Bellman, a scalable algorithm that improves performance and transfer speed over reward regression based SF baselines and standard baseline algorithms in continual-learning and goal-conditioned control tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.