acceptodds
Under review as a conference paper at ICLR 2027

Control-centric Representation Learning using Action-free Datasets with Distinct Policies

Abstract

Action- and reward-free data are available in settings such as robot manipulation, video game play, and human egocentric recordings. Prior work uses these data to learn visual representations for downstream control with limited action-labeled data. However, time-correlated visual distractions can degrade the representations learned by these methods. Critically, in many common visual control tasks, a naive encoder cannot distinguish noise from controllable features, leading to poor generalization. In this work, we consider action-free trajectories grouped by known data-collection policy within a shared environment. We introduce CARDPol, a representation-learning objective that predicts policy identity from pairs of observations. Under appropriate assumptions, we show that the ground-truth endogenous representation minimizes our objective. Empirically, our proposed loss learns representations that exhibit reduced sensitivity to time-correlated exogenous noise. CARDPol improves distracted LIBERO success by 3–11 percentage points and DMC video-distraction returns by 25–53% over the strongest evaluated baseline in each setting. On human egocentric video from EgoDex, it achieves the lowest mean wrist-position decoding error, with the closest baseline exhibiting 15% worse error. Our results demonstrate that policy identity can provide useful supervision for learning control-relevant representations from high-dimensional observations without action labels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.