Near to Far Self Distillation for Offline Goal-Conditioned Reinforcement Learning
Abstract
Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals from a fixed dataset. Existing methods use value estimates or subgoals to guide policy learning, but supervision for current decisions can remain weak when goals are distant. Inspired by the dense predictive supervision of on-policy distillation (OPD), we use teacher predictions as direct learning targets. This raises two challenges in offline GCRL: (1) the Guidance Selection Challenge, identifying privileged information that provides useful guidance toward the distant goal; and (2) the Information Gap Challenge, transferring such guidance to a student without access to the privileged information. To address these challenges, we propose NEFD, a Near-to-Far Self-Distillation method for offline GCRL. We observe that near-future states reveal local outcomes of recorded behavior and thus provide informative privileged references for current decisions. NEFD uses these states as teacher-only information and weights teacher guidance by estimated goal progress. To bridge the information gap, NEFD performs action and representation self-distillation within a shared actor: the former matches teacher action distributions, while the latter uses teacher decision representations as intermediate targets. Experiments on 29 OGBench datasets show that NEFD improves policy learning when distant goals provide weak supervision for current decisions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.