acceptodds
Under review as a conference paper at ICLR 2027

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

Abstract

Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks through self-supervision. However, under failure termination, established CRL considers pre-failure future goals only when constructing positive samples, discarding the occupancy mass absorbed by failure. This omission induces a systematic overestimation bias in the implied goal-reaching values, giving short surviving futures disproportionately strong supervision. Such distortion can reinforce failure-prone behaviour through catastrophic failure bootstrapping or favour unsustainable goal reaching. To address this problem, we introduce two simple yet theoretically rigorous corrections: mass-weighted InfoNCE corrects the overweighting of short surviving futures in critic learning, and a log-survival-mass score restores the missing mass in policy optimization under the population formulation. The resulting method, Safe Contrastive Reinforcement Learning (Safe-CRL), uses the one-bit failure termination signal as its only safety supervision. Across twelve failure-prone robot navigation and locomotion tasks, Safe-CRL consistently improves mean survival time and achieves higher mean time at goal than Scaling-CRL, with substantial gains in nine tasks. It also preserves depth scalability, with deep policies exhibiting complex failure-avoidance behaviours. These results establish survival-mass correction as an effective completion of CRL for failure-prone goal-conditioned control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.