acceptodds
Under review as a conference paper at ICLR 2027

SHARV: A Simple Multi-Horizon Approach to Representation and Value Learning in Reinforcement Learning

Abstract

Effective use of collected experience is essential for sample-efficient reinforcement learning. However, existing methods often rely on fixed horizon configurations, limiting their ability to exploit learning signals across environments with different temporal characteristics. We show that short- and long-horizon prediction have environment-dependent effects and that exhaustively combining prediction targets does not consistently improve performance. We present SHARV, a simple approach that exploits multiple horizons for representation and value learning. Our representation learning combines independently learned short- and long-horizon representations to support policy and value learning across diverse environments. For value learning, we characterize internal TD residual components not determined by the first-start target error, motivating Multi-Start Bootstrapping to control these residuals through supervision across multiple starting points. SHARV achieves the highest mean aggregate score among strong baselines on each of four benchmark suites. Ablation studies support the contribution of both components, and observed reductions in internal TD residuals provide empirical support for our theoretical analysis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.