acceptodds
Under review as a conference paper at ICLR 2027

Bellman Demand and Critic Response: A Spectral View of Continual RL

Abstract

Continual reinforcement learning requires agents to retain the ability to learn as tasks change over time, yet plasticity loss in neural networks can impede learning on subsequent tasks. In SAC, policy improvement relies on value estimates provided by the critic; therefore, whether the critic can learn the value function of a new task within a finite update budget is critical to continual adaptation. We study this problem through the discrepancy between the Bellman target and the current prediction, examining how the value corrections required by a new task align with the critic's learning responses across different directions. When the dominant correction demand of a new task is concentrated in directions to which the critic responds slowly, this mismatch can hinder value correction. To quantify this difficulty, we derive a demand-weighted spectral burden that combines the fraction of demand carried by each direction with its relative response speed, characterizing the relative directional difficulty of correcting the current residual. Motivated by this analysis, we adopt Periodic Spectral Projection (PSP) in continual SAC, constraining the singular values of each linear layer in both critics to a prescribed interval to modulate their learning responses at the weight level. Experiments on Meta-World and CompoSuite show improved subsequent-task acquisition with PSP. The spectral analyses characterize how the current Bellman correction demand is distributed across the critic's response directions during adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.