acceptodds
Under review as a conference paper at ICLR 2027

Don’t Relearn What You Already Know: using Residual Reward Learning for RLHF

Abstract

Human preferences are expensive, yet experts can often specify some of the components of a desired reward programmatically. Recent work has shown that such proxy rewards can be combined with preference learning by learning only their residual. But why does residual reward learning outperform simply adding a proxy to a conventionally learned reward model, and when should we expect it to do so? We identify two mechanisms. First, residual learning preserves the decomposition of the true reward, whereas naive addition double-counts components already captured by the proxy. Second, when the proxy explains sufficient variation in the true reward, the residual constitutes a lower-variance and therefore simpler learning target. We theoretically characterize these conditions and test them through controlled experiments across eight RL environments. Finally, we demonstrate that the same principle extends to LLM alignment by combining a programmatic proxy with human preferences for summarization. Our results characterize when prior reward knowledge can reduce the burden placed on preference learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.