acceptodds
Under review as a conference paper at ICLR 2027

DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction

Abstract

While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the perception gap, where static visual inputs lack temporal motion cues; the latency gap, where inference delays render actions obsolete; and the control gap, caused by the open-loop action chunk execution without real-time adjustment. In this work, we propose DSDyn-VLA, a Slow-Fast Dual-Stream Dynamic manipulation framework that integrates motion-aware foresighted planning with real-time residual correction. The slow Flow-Planner serves as a macro-planner. By enhancing the VLA with optical flow for temporal perception and a future state awareness mechanism to preemptively offset inference latency, it produces globally consistent, motion-aware action chunks. Complementing this, the fast Res-Refiner employs a lightweight RL policy to inject high-frequency, closed-loop corrections into the planned action chunks based on real-time observations. In addition, we introduce DynBench, a MuJoCo-based benchmark for dynamic object manipulation that comprises nine tasks. Extensive experiments demonstrate that DSDyn-VLA reduces the failure rate by over 76% compared to current SOTA method in high-latency setting on the Kinetix dynamic benchmark, while achieving about 6 the success rate of PI0.5 in real-world dynamic settings and about 5 on DynBench. We will open-source all the code and weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.