acceptodds
Under review as a conference paper at ICLR 2027

Learning to Act under Visual Interruptions with Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which trains VLA policies to handle missing visual inputs and temporarily supplies predicted images when cameras fail. It uses optical-flow extrapolation for missing wrist views or all-vision loss, and an action-conditioned world model for missing third-person views, withdrawing the predictions when they become unreliable. Experiments on and GR00T N1.5 show that MINT improves task success under camera loss over the original models. Experiments on the AgiBot G2 further demonstrate real-robot closed-loop deployment under camera loss.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.