Bridging Reward Regression and Projected Bellman Equation with Successor Features
Abstract
Successor features (SF) decompose the Q-function into a successor feature map and a task vector, enabling efficient transfer in continual and goal-conditioned reinforcement learning. The Q-function estimate depends critically on how the task vector is estimated. We formalize the two estimators used in the literature—reward regression and the projected Bellman equation—prove that they generally yield different solutions, and characterize when they coincide. We unify them through an interpolation parameter and establish a stability–approximation tradeoff: reward regression stabilizes off-policy updates, while moving toward the solution of projected Bellman equation tightens the worst-case approximation guarantee. Building on this, we develop SF-Bellman, a scalable algorithm that improves performance and transfer speed over reward regression based SF baselines and standard baseline algorithms in continual-learning and goal-conditioned control tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.