RESTA: Time-Scaling Augmentation for Data-Efficient Robot Imitation Learning
Abstract
Robot manipulation policies trained through imitation learning commonly rely on costly demonstrations recorded in one task direction and at one temporal resolution. We propose **Re**versal and **S**tride-based **T**rajectory **A**ugmentation (RESTA), which exploits time-scaling data augmentation to improve data efficiency through trajectory reversal and temporal resampling. Trajectory-reversal augmentation reverses and relabels demonstrations between paired tasks with opposite objectives, reusing recorded rollouts without additional environment interaction. Because contact-rich manipulation is not exactly reversible, we treat reversal as an approximate paired-task prior rather than a physical symmetry. Temporal-stride augmentation varies observation and action sampling strides and conditions the policy on the resulting offsets, enabling fixed or randomized temporal schedules. Experiments across 16 paired bimanual manipulation tasks show that trajectory reversal is particularly effective in the low-data regime and that temporal-stride augmentation improves success under randomized temporal schedules. With 25 and 50 independently collected demonstrations per task, RESTA achieves 78.4% and 85.2% average success, respectively, matching or exceeding baselines trained with twice as many demonstrations (77.3% and 85.4%). These results show that RESTA can improve data efficiency in robot imitation learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.