Temporal Contrastive Test-Time Training for Robotic Manipulation
Abstract
Robot memory faces two directions in time: the past determines what can be known now, while the future reveals what was worth knowing. Test-time training (TTT) gives this asymmetry an optimization form: the inner objective determines what writes into memory, while outer task losses determine which updates are useful. However, common TTT objectives derive self-supervision primarily from the current observation, providing no direct objective for modeling relations across observed history. We introduce Temporal Contrastive Test-Time Training (TC–TTT), a contrastive inner objective that augments reconstruction by distinguishing the current representation from historical alternatives conditioned on an earlier observation. We further study how to select historical observations for temporal comparison, examining how sampling stride and candidate count affect policy performance. In a simplified model, TC–TTT adds a velocity-like term to the reconstruction gradient, making first-order temporal change explicit. Across nine RMBench tasks, TC–TTT improves mean success from 31.8% to 51.8% over reconstruction-based TTT under the same 77M-parameter Diffusion Policy and action-gradient window. These results identify the inner objective as an important design axis for online robot memory and show that historical contrast can substantially improve control without extending the action-gradient horizon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.