STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning
Abstract
Real-world robot learning increasingly relies on heterogeneous data, but demonstrations and rollouts often mix useful progress with stalls, corrections, and suboptimal behavior. Effective policy learning therefore requires frame-level advantages that distinguish reliable local progress from failures and regressions. We propose **S**elf-supervised **T**emporal **E**nsemble **A**dvantage **M**odeling (**STEAM**), a label-free method that learns such advantages from expert demonstrations. STEAM trains an ensemble of temporal-offset predictors on frame pairs within expert trajectories, using the normalized temporal offset between two frames as a self-supervised signal. Each predictor maps a frame pair to a distribution over temporal offsets, which is converted into a scalar advantage. STEAM then takes the minimum advantage across the ensemble to score mixed-quality rollout data conservatively. Across real-world bimanual towel folding, chip checkout, table clearing, and cola restocking tasks, STEAM identifies stalls, failures, and recoveries. When combined with CFGRL, STEAM achieves the highest policy success rate on all four tasks, averaging 83.5% against 58.8% for the best baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.