Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models
Abstract
Vision-Language-Action (VLA) models achieve remarkable flexibility and generalization beyond classical control paradigms. However, most prevailing VLAs are trained under a single-frame observation paradigm, which leaves them structurally blind to temporal dynamics. Consequently, these models degrade severely in non-stationary scenarios, even when trained or fine-tuned on dynamic datasets. Existing approaches either require expensive retraining, or suffer from latency bottlenecks and poor temporal consistency across action chunks. We propose Pace-and-Path Correction (PPC), a training-free, closed-form inference-time operator that wraps any chunked-action VLA. Under a local motion model, joint minimization of a single quadratic cost yields an interior solution with two orthogonal channels. The pace channel compresses execution along the planned direction, while the path channel applies an orthogonal spatial offset, jointly absorbing the perceived dynamics within the chunk window. We evaluate our approach on a comprehensive diagnostic benchmark MoveBench designed to isolate motion as the sole controlled variable. Under the controlled MoveBench protocol, PPC outperforms the evaluated training-free wrappers and dynamic-adaptive baselines, raising success by up to 28.8% and 25.9% in absolute terms over foundational VLAs in dynamic-only and static-dynamic mixed environments, respectively. We further validate PPC on a real robot with a backbone fine-tuned only on static demonstrations, increasing dynamic grasping success from 0% to 68.3%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.