acceptodds
Under review as a conference paper at ICLR 2027

Successor-Error Amplification in Temporal-Difference Learning

Abstract

With function approximation, a temporal-difference (TD) update at one state can change the prediction at its successor, an effect we call *successor shift*. The successor is special because its prediction also appears in the bootstrap target used to form the update. We show that this creates a feedback mechanism in which the successor's value error can amplify itself through the TD update, which we call *successor-error amplification*. Under simple conditions on deterministic chains, this amplification is strong enough that the TD update increases the successor's squared value error in expectation. To mitigate this amplification, we derive Successor-Shift-Controlled TD (SSC-TD), which suppresses successor shift while preserving the norm of the TD update. An adaptive version estimates the shift's average effect on the successor's squared value error, and activates suppression when the estimate indicates harm. Experiments show adaptive SSC-TD lowers lifetime MSE relative to TD on a majority of Atari environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.