Efficient Online Estimation of Achievable Goal Distributions During Training for Goal-Conditioned Reinforcement Learning
Abstract
Goal-conditioned reinforcement learning (GCRL) trains a single policy to achieve diverse goals. Central to GCRL is the achievable goal distribution , the distribution of goals that the current policy can achieve. Accurate estimation of helps curriculum learning methods in GCRL sample behavioral goals (the goals used to collect training data) of intermediate difficulty, which substantially improves training efficiency. Estimating online, however, is challenging: the distribution evolves throughout training, and the estimator must track this moving target with data that are both fresh and sufficient. Existing methods meet only one of these requirements: replay-buffer-based estimation requires no extra environment interaction but is biased toward where past policies happened to visit, whereas policy evaluation faithfully reflects the current policy but demands a large number of fresh environment samples for every estimate. We empirically find that evolves smoothly during training, so historical policy evaluation data remain informative for estimating the current . Exploiting this finding, we propose PE-GCRL (GCRL with **P**eriodic Policy **E**valuation), which interleaves training with a small number of policy evaluations and estimates by reusing the accumulated historical evaluation data. Experiments on four mainstream GCRL tasks show that, across multiple curriculum algorithms, PE-GCRL reduces the estimation error of by roughly an order of magnitude, keeps behavioral goals at the frontier of achievability, and improves learning speed and final success rates with less than 5% training-time overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.