Relative Deadline Regret for Goal-Conditioned Reinforcement Learning
Abstract
Goal-conditioned reinforcement learning trains a policy to reach a desired goal from the current state with a success signal provided only upon arrival in the sparse reward setting.Policy learning near obstacles is difficult because actions that begin useful detours can receive the same immediate reward as actions that make no progress.Bellman updates compare these actions through estimated future rewards whose errors can affect policy optimization.Relative Deadline Regret (RDR) is proposed to compare candidate actions through completion probability, the probability of first arrival within a step limit called a deadline.Regret measures the completion probability lost against the best candidate as a share of the positive probability of that candidate.An arrival predictor learns from observed arrivals and unfinished attempts to compare candidate actions without Bellman action value targets.The policy fits a weighted average that favors smaller regrets across deadlines and follows a local direction of higher predicted completion called the completion gradient.Theoretical analysis bounds candidate and executed action losses and establishes convergence under stated stability conditions for the expected policy fitting update.Experimental results show the highest success rate averaged equally over the evaluated manipulation and navigation tasks and connect the candidate comparisons to learned actions near obstacles.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.