acceptodds
Under review as a conference paper at ICLR 2027

In-Context Test Time Training for Adaptive Robot Policy Learning

Abstract

During deployment, robots continually receive new observations that can inform subsequent actions. Yet many robot policies operate with fixed parameters and a short observation history, limiting their use of accumulated experience. Learning from this experience requires both selecting useful information and training the policy to act with a memory that changes over time. We introduce In-Context Test Time Training (In-Context TTT), a framework that combines adaptive fast-weight memory with selective writing for robot policy learning. TTT layers encode incoming observations into fast weights, while a learned gate controls the strength of each memory update. To train this recurrent adaptation process, we develop Trajectory-Consistent Adaptation (TCA), which supervises action prediction along the fast-weight trajectories generated by ordered demonstrations. After chunk-level initialization, TCA carries fast weights across an entire episode, accumulates supervised gradients at successive adaptation states, and updates the slow parameters once at the episode boundary. This aligns the memory evolution used in training with deployment, where fast weights change while slow parameters remain fixed. On the six longest RoboCasa GR-1 tasks, selected by an average demonstration length of at least 500 steps, In-Context TTT achieves 53.3% mean success, compared with 50.0% for No-update SFT and 47.0% for the original policy. Real-robot evaluations obtain 7/10 successes in banana picking, 6/10 in cup stacking, and 4/10 in arranging digit-shaped blocks. Together, selective writing and trajectory-consistent supervision provide a way to learn from interaction history using demonstration data. Our core implementation is available at https://anonymous.4open.science/r/In-Context-TTT-E538.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.