acceptodds
Under review as a conference paper at ICLR 2027

The Horizon of Reward Guidance

Abstract

Preference feedback can leave parts of the reward function unidentified even when it is sufficient to support a useful policy update. As the policy changes, previously irrelevant ambiguity can become important, while the deployed reward can adapt only finitely fast. We study how long existing preference information and finite reward adaptation can continue to certify positive policy progress, and call this duration the support horizon. In a scalar model, we characterize the support horizon exactly and prove that pointwise feasibility can persist after every admissible reward trajectory has lost the certificate. We show that the same separation can arise endogenously in finite policy dynamics and locally under ordinary Euclidean policy gradient. The resulting dynamic boundary also determines which preference uncertainty must be resolved to extend guidance. Under Bradley–Terry feedback, the number of additional labels required scales with the inverse square of the distance to the dynamic feasibility boundary, up to logarithmic dependence on the confidence level. These results connect reward guidance reachability with targeted preference acquisition and show that continued guidance is jointly limited by reward uncertainty and adaptation speed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.