Don’t Relearn What You Already Know: using Residual Reward Learning for RLHF
Abstract
Human preferences are expensive, yet experts can often specify some of the components of a desired reward programmatically. Recent work has shown that such proxy rewards can be combined with preference learning by learning only their residual. But why does residual reward learning outperform simply adding a proxy to a conventionally learned reward model, and when should we expect it to do so? We identify two mechanisms. First, residual learning preserves the decomposition of the true reward, whereas naive addition double-counts components already captured by the proxy. Second, when the proxy explains sufficient variation in the true reward, the residual constitutes a lower-variance and therefore simpler learning target. We theoretically characterize these conditions and test them through controlled experiments across eight RL environments. Finally, we demonstrate that the same principle extends to LLM alignment by combining a programmatic proxy with human preferences for summarization. Our results characterize when prior reward knowledge can reduce the burden placed on preference learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.