Relative Temporal Alignment for Real-Time Execution of Vision-Language-Action Policies
Abstract
Vision-Language-Action (VLA) policies commonly leverage action chunking to amortize the high computational cost of model inference. Action chunking reduces amortized computation but also lowers the rate at which fresh observations enter the policy. Synchronous execution may pause the robot between chunks, whereas asynchronous execution maintains continuous control but hands over a chunk conditioned on an earlier observation. Existing approaches address this challenge through chunk-boundary smoothing, future robot-state prediction, or stepwise visual feedback, yet they do not explicitly relate robot–environment changes during asynchronous inference to action residuals (i.e., corrections to the pending action sequence), limiting low-cost temporal alignment of pending actions. We propose Relative Temporal Alignment (RTA), a lightweight correction method decoupled from the base VLA policy. RTA keeps the base policy frozen and uses relative robot–environment changes during inference to efficiently correct each incoming action chunk before execution. This aligns actions conditioned on the observation at inference onset with the robot–environment state available when inference completes. Experiments across static simulation, dynamic simulation, and real-robot platforms show that RTA maintains robust performance under varying inference delays and improves task completion rates while incurring minimal correction overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.