BLUR: Bootstrapped Latent Utility Reward Learning from Trajectory-Level Performance
Abstract
In complex reinforcement learning tasks, handcrafted dense step-wise rewards can substantially improve exploration efficiency and policy performance. However, their construction requires extensive domain knowledge and scales poorly to new tasks. Existing automatic reward-learning methods reduce the need for manual intervention, but often struggle to model high-dimensional sequential behaviors and suffer from severe cold-start problems when successful trajectories are scarce. To address these limitations, we propose Bootstrapped Latent Utility Reward Learning (BLUR), which learns context-dependent dense rewards solely from trajectory-level performance signals. BLUR first constructs relative utility supervision by comparing trajectory performance against the recent policy level, guiding the policy from random failures toward progressively better failures and eventually successful behaviors. The temporary relative model is then discarded, and a new reward model with fixed absolute semantics is trained from both successful and unsuccessful trajectories. Finally, this model is frozen and used to optimize an independently initialized target policy. The reward model employs an LSTM to encode temporal behavioral context and VaDE to model latent behavioral structures, from which step-wise rewards are generated using trajectory-performance statistics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.