acceptodds
Under review as a conference paper at ICLR 2027

Transferring Relative Motion Into Action for Dynamic-Centric Manipulation

Abstract

Vision-language-action (VLA) models perform well in static manipulation but remain challenged by independently moving targets. As objects continue to move during inference, generated actions can become misaligned with the scene at execution time. Existing approaches address this perception-to-execution gap through execution scheduling or scene modeling, but often leave observed target dynamics weakly coupled to action generation. We propose Motion Into Action (MIA), a framework that couples object-centric perception with dynamics-conditioned action generation. MIA maps task-relevant optical-flow history to a motion-induced latent action that parameterizes a motion-aligned Gaussian source for diffusion, guiding the action head toward executable actions in fewer denoising steps. Context-aware asynchronous execution further overlaps inference and execution while promoting temporal consistency across action chunks. Across five DOMINO tasks, achieves an average success rate of 41.6%, relatively outperforming the strongest baseline by 55.22%. On two tasks, it maintains performance when denoising is reduced from 20 to 5 steps. Across five real-world tasks, MIA achieves an average success rate of 88%, with qualitative examples illustrating transfer to faster motion, occlusion, and moving distractors beyond the demonstration conditions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.