acceptodds
Under review as a conference paper at ICLR 2027

MEF-VLA: Motion-to-Event Foresight for Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models have advanced robot manipulation, and predictive world modeling further supports action learning through future-state prediction. Despite these advances, policies can still deviate from the instructed sequence or fail to complete individual interactions in multi-stage tasks. To improve the guidance provided by visual foresight, we revisit which future to predict and what to represent. A fixed prediction horizon can end during intermediate motion or extend beyond the next interaction, resulting in inconsistent guidance for task progress. Even with a reference aligned with task progress, complete future-frame targets retain static scene appearance while leaving the spatial change from the current state implicit. We introduce MEF-VLA (Motion-to-Event Foresight), which learns the motion remaining until the next interaction event. Event frames marking upcoming interactions define adaptive prediction horizons. Remaining event-flow represents the motion toward these shared references as state-dependent spatial displacement. Learned event queries predict remaining-motion features and condition the action expert, adding negligible inference overhead. In simulation, MEF-VLA improves success rates over the vanilla VLA by 1.6 percentage points on LIBERO, 8.9 on zero-shot LIBERO-Plus, and 9.0 on five-task sequences in CALVIN DD. Across three real-world task families, MEF-VLA trained on one- and two-stage instructions achieves 58.9% average success on held-out three-stage instructions, compared with 34.4% for the vanilla VLA. Ablations further demonstrate the benefits of both event-aligned horizons and remaining-motion supervision for task success and zero-shot robustness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.