Under review as a conference paper at ICLR 2027
Finite-Horizon Hamiltonians for Value Geometry in Goal-Conditioned RL
Abstract
In goal-conditioned reinforcement learning (GCRL), quasimetric learning models goal-reaching costs as quasimetric distances, connecting local constraints to global value geometry. Its local constraints, however, should reflect the direction-dependent effects of control composition over a finite horizon together with environmental feasibility. We propose *ReQRL*, which constrains the critic's value gradients through finite-horizon reachability. Drawing on state-constrained optimal control, we decouple dynamical reachability from boundary geometry, estimating both from data. On *OGBench*, our method outperforms or rivals existing quasimetric approaches and other offline GCRL methods.
open until 14 Dec 2026
est. 32% chance this paper gets accepted at ICLR 2027.
Reject 68%Accept 32%
What do you think this paper will get?
All positions stay anonymous.
Related papers
Loading the map…
Discussion (0)
Sign in to comment.