Transferring Relative Motion Into Action for Dynamic-Centric Manipulation
Abstract
Vision-language-action (VLA) models perform well in static manipulation but remain challenged by independently moving targets. As objects continue to move during inference, generated actions can become misaligned with the scene at execution time. Existing approaches address this perception-to-execution gap through execution scheduling or scene modeling, but often leave observed target dynamics weakly coupled to action generation. We propose Motion Into Action (MIA), a framework that couples object-centric perception with dynamics-conditioned action generation. MIA maps task-relevant optical-flow history to a motion-induced latent action that parameterizes a motion-aligned Gaussian source for diffusion, guiding the action head toward executable actions in fewer denoising steps. Context-aware asynchronous execution further overlaps inference and execution while promoting temporal consistency across action chunks. Across five DOMINO tasks, achieves an average success rate of 41.6%, relatively outperforming the strongest baseline by 55.22%. On two tasks, it maintains performance when denoising is reduced from 20 to 5 steps. Across five real-world tasks, MIA achieves an average success rate of 88%, with qualitative examples illustrating transfer to faster motion, occlusion, and moving distractors beyond the demonstration conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.