T2Align: Tempo and Transition Alignment for Asynchronous Robot Manipulation
Abstract
Chunked imitation learning policies, especially vision-language-action (VLA) models, have shown impressive performance in real-world robot manipulation by predicting temporally extended action chunks. However, deploying chunked policies in time-critical dynamic tasks remains challenging: task states change during execution, while success depends on reacting to these changes and completing motions within narrow temporal windows. We present Align, a robot manipulation system that provides hierarchical tempo and transition alignment for asynchronous execution of chunked policies. At the execution level, Align preserves the policy-intended execution tempo under safety constraints. At the action chunk level, Align aligns returned chunks to the current execution timeline and stabilizes their transition through continuity-consistency blending and similarity-gated transition control. Align is plug-and-play with off-the-shelf policies and improves their applicability to time-critical dynamic tasks. Experiments with widely used policies show that Align substantially improves success rate and stability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.