acceptodds
Under review as a conference paper at ICLR 2027

Not All Transitions Are Equal: Uncertainty-Driven Contrastive Learning for Visual RL

Abstract

Recent works show that prioritizing transitions from the replay buffer can improve representation learning by focusing on informative experiences, potentially guiding exploration and improving sample efficiency. We present UCRL: **U**ncertainty-Driven **C**ontrastive Learning for Visual **R**einforcement **L**earning, a method that leverages transition uncertainty to prioritize experiences for contrastive representation learning. Unlike existing contrastive RL methods, which uniformly sample transitions for representation learning, UCRL modifies only the transition sampling distribution of the contrastive objective. It prioritizes transitions according to epistemic uncertainty estimated from an ensemble of critics. We theoretically establish that nonzero epistemic uncertainty lower-bounds the encoder gradient induced by the critic objective, identifying epistemic uncertainty as an observable proxy for representation learning potential. We apply UCRL to CURL and evaluate the resulting method on the DeepMind Control Suite 100K and Atari 100K benchmarks, achieving relative performance gains of 9.5% and 14.8%, respectively, over CURL. Across five of the six main DeepMind Control Suite tasks, UCRL reaches CURL's 100k performance using 46.4% fewer environment interactions on average, with reductions of up to 65% across individual tasks. Code will be available upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.