Data That Selects Itself: Scalable Self-Curation for Robot Learning
Abstract
Robot imitation learning relies on demonstrations of varying quality, making effective data curation essential for scalable and data-efficient policy learning. However, learning a useful curation signal remains challenging when trusted reference demonstrations and policy-rollout feedback are unavailable. We present Trajectory Progress Efficiency Consensus (TPEC), a self-supervised framework that learns to rank demonstrations from small, uncurated reference sets sampled directly from the candidate dataset. Using lightweight progress predictors, TPEC aggregates whole-trajectory efficiency scores into a consensus ranking for action-chunk selection, enabling data curation without curated references or online policy rollouts. Experiments demonstrate the effectiveness of TPEC across demonstration-quality regimes and policy scales. On RoboTwin 2.0 and in real-robot experiments, TPEC matches or exceeds full-data joint fine-tuning of the billion-parameter VLA model π0 using only 50% of the available action chunks. Furthermore, TPEC scales to large-scale pretraining, highlighting its potential for practical robot learning pipelines. Enabled by its in-dataset selection strategy, TPEC achieves curation-preparation speedups of 2,257× over the gradient-based CUPID and 4.53× over the progress-reward-based WARP-RM in VLA data curation. Code will be public.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.