acceptodds
Under review as a conference paper at ICLR 2027

Relative Bellman Constraints for Value Learning in Residual Reinforcement Learning

Abstract

Residual reinforcement learning (RL) adapts pretrained robot policies through bounded action corrections, making it appealing for real-robot learning where exploration and intervention are costly. However, value learning in residual RL typically focuses on absolute returns, whereas adaptation with a frozen base policy depends on resolving the value variations induced by these corrections. This motivates making value differences an explicit focus of supervision over the replay distribution, rather than learning them only indirectly from pointwise value targets. In this paper, we propose **ReVRA** (Relative Value Learning for Residual Adaptation), a plug-and-play value-learning module that explicitly supervises critic-induced value differences. ReVRA incorporates relative Bellman constraints alongside standard TD learning, using paired replay transitions to supervise value differences within the existing critic. Across simulated locomotion, manipulation, and VLA-based adaptation, ReVRA consistently improves residual RL performance. On real-robot manipulation tasks, it achieves higher task success with fewer interventions under the same limited interaction budget. The project code is available at https://anonymous.4open.science/r/ReVRA-DF70.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.