acceptodds
Under review as a conference paper at ICLR 2027

AQRL: Quasimetric Reinforcement Learning with Split State Encoding and Approximate Nearest Completion

Abstract

Goal-conditioned tasks require an agent to reach a state satisfying a given goal. Such a goal often constrains only some state coordinates, which we define as specified coordinates; the rest are free coordinates. Quasimetric Reinforcement Learning (QRL) is a goal-conditioned reinforcement learning method that performs well. It takes states as inputs to estimate state-reaching cost. However, it has two limitations. First, its actor receives sampled real states as goals during training but artificially constructed states during testing, creating a mismatch in goal inputs. Second, QRL trains its actor to reach a single specific state for each training sample, even though other states that satisfy the same goal may have lower reaching costs. To address both limitations, we propose Approximate nearest-goal Quasimetric Reinforcement Learning (AQRL) with two components: split state encoding (SSE) and approximate nearest completion (ANC). SSE encodes specified and free coordinates separately, enabling the actor to use only specified-coordinate latents as goal inputs in training and testing, resolving the mismatch. ANC searches over free-coordinate latents during training to find a completion with lower estimated reaching cost. Across fourteen online tasks with five training seeds, AQRL outperforms QRL on thirteen tasks and ranks first among all compared methods on twelve, achieving 66.1% mean success versus 24.5% for QRL and 45.3% for the strongest compared baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.