acceptodds
Under review as a conference paper at ICLR 2027

The Reward Horizon Problem: Learning From Delayed Feedback

Abstract

People increasingly use language models to pursue objectives that are realized over long time horizons. However, post-training rewards are typically computed over short horizons, which creates a mismatch between the reward signal and the ultimate objective. For example, while a short-horizon training signal may reward code that runs, does it reward code that stays stable? We study this _reward horizon problem_ in a software engineering setting, where the consequences of a code change may not be observed for several days or months. Using code corrections as a source of delayed feedback, we find substantial differences between the signals available at short and long horizons. Short-horizon feedback omits information about long-horizon objectives and encodes a different distribution of label sources. We then analyze how critic models trained on long-horizon outcomes may distill some of this delayed feedback into short-horizon rewards.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.