acceptodds
Under review as a conference paper at ICLR 2027

DreamIF: Intention-Guided Sparse Dynamics Prediction for Long-Horizon Vision-Language-Action Manipulation

Abstract

Leveraging robust semantic representations of pre-trained VLMs, Vision-Language-Action (VLA) models build end-to-end multimodal perception-to-action mappings to boost semantic alignment and task generalization for embodied manipulation. Nevertheless, existing VLA paradigms lack task-oriented dynamics inference, making them vulnerable to irrelevant background redundancy and cumulative errors in long-horizon manipulation. This work proposes Dream Intention Flow (DreamIF), a novel intention-guided embodied foundation model that enables VLA to focus on intention-relevant core cues and mitigate environmental redundancy and long-horizon errors. We design an intention-driven dynamics inference mechanism: optical-flow-guided motion saliency locates Intention-Associated Dynamic Visual Regions (IADVR) to filter redundant backgrounds, capture physical transition rules, and extract action intentions. The model synchronously predicts multi-frame future IADVR features and embeds these temporal priors into the native action representations, strengthening temporal modeling and shifting control from reactive to predictive. Evaluated on the comprehensive LIBERO-PLUS benchmark under matched training data, DreamIF attains 85.7% on average, and extending its method to a backbone lifts that backbone's average success rate by 15.0 points. Notably, on the long-horizon LIBERO-PLUS Long suite it achieves an 85.5% success rate, 8.9 points above the matched baseline, and on eight real-world robotic tasks involving environmental noise and long-horizon manipulation it raises the average success rate by 15.6 points. Experimental results validate that DreamIF greatly enhances the ability to model task-oriented dynamics and remedies the inherent limitations of conventional VLA models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.