Acting at the Moment of Consequence: Physical Foresight for VLA Action Refinement
Abstract
Vision-language-action (VLA) models can produce coherent manipulation behavior yet fail during physical interactions that require precise control. Refining their actions requires identifying when local adjustments can still affect an operation’s outcome. We introduce PREEMPT, a framework for selectively refining the actions of a fixed VLA policy. A V-JEPA-based anticipator foresees physical decision windows and initiates local correction. A latent world model predicts short candidate trajectories, and a completion-time evaluator relates these forecasts to the likelihood and timing of operation completion. These estimates guide smooth corrections constrained by the policy’s intended motion, while predicted trajectory diversity guides further search allocation. We evaluate PREEMPT on LIBERO and RoboCasa with , on LIBERO with OpenVLA-OFT, and on four real-robot tasks with an Agilex dual-arm robot. With , PREEMPT increases success from 96.7% to 97.1% on LIBERO and from 64.9% to 68.6% on RoboCasa. Real-robot success rises from 72.5% to 82.5%, a gain of 10 percentage points. In a LIBERO-10 scheduling comparison, window-based search uses 58.4% fewer candidate evaluations than searching at every replan.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.