The World Does Not Pause: Real-Time Benchmarks and Training-Free Acceleration for Dynamic Manipulation
Abstract
Dynamic manipulation requires timely actions as the world continues to evolve during policy inference. In vision-language-action (VLA) and world action models (WAM), inference latency can leave actions conditioned on outdated observations. We introduce Dynamic-LIBERO and Dynamic-RoboTwin, which combine controlled target motion with inference-time world evolution while largely preserving established task goals for frozen-policy evaluation. We propose PACE (Policy Acceleration through Cached Execution), a training-free runtime that reuses observation representations and action computation across replanning calls and integration steps, and compiles the remaining computation without policy fine-tuning. Across two simulated benchmarks and two policies, PACE accelerates inference by up to 3.8× and improves dynamic task success by up to 42.8 percentage points over standard deployments, achieving the highest success in all eight dynamic settings. On a physical robot, PACE reduces per-chunk inference stalls from 159 to 69 ms, maintains or improves static-task success, shortens completion times, and increases dynamic success counts. Together, these results highlight that reducing inference latency is critical for successful dynamic manipulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.