Fast Thinking, Steady Control: Coherent Action-Chunk Transitions for Real-Time VLA Inference
Abstract
Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. To improve responsiveness to the environment, a growing body of work has focused on accelerating VLA inference. However, in real-world deployment, faster VLA inference does not necessarily translate into better control. Under continuous asynchronous inference, lower inference latency increases the frequency of chunk switching, which introduces inter-chunk discontinuities and latency-alignment errors, leading to discontinuous, jerky motion and even degraded task success. In this paper, we propose F2S-VLA, a training-free execution framework that translates fast asynchronous VLA inference into steady and coherent robot control. To mitigate inter-chunk discontinuity, we propose Noise-Decoupled Chunk Prediction, which generates two action chunks in a single forward pass with reference and newly sampled noise to separate observation-driven trajectory updates from noise-induced variations, and Anchor-Guided Trajectory Fusion, which uses the ongoing trajectory as an anchor to selectively fuse the new predictions, adaptively balancing responsiveness and continuity without task-specific thresholds. To eliminate latency-alignment errors, we introduce Latency-Aligned Action Resampling, which registers action chunks on a shared physical timeline and resamples them at the measured takeover time using the controller's native interpolation rule, improving temporal alignment granularity from tens of milliseconds to less than two milliseconds with negligible overhead. Experiments demonstrate that F2S-VLA outperforms RTC by 9.0% on RoboCasa365 benchmark and achieves a 93.9% average success rate across three long-horizon real-robot tasks. Our code is available on Anonymous Github.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.