acceptodds
Under review as a conference paper at ICLR 2027

FlowLap: Execution-Aware Inference for Flow-Matching Video-Action World Models

Abstract

Flow-matching video-action models jointly generate visual futures and action chunks, but rollout latency can stall robot execution at chunk boundaries. Although asynchronous action-chunk methods such as Real-Time Chunking (RTC) can hide part of this latency, directly adapting RTC to flow-matching video-action models that commit real context or KV cache only at chunk boundaries can produce stale-context action candidates; VLASH, LingBot-VA 2.0, and DexWorldModel require model-specific training or pretraining to preserve success under overlap. We introduce FlowLap, an execution-aware inference framework with two components: a speculative joint rollout that overlaps inference with action execution while preserving the committed action prefix, and execution-state calibration that conditions the action branch on the state at which the candidate will run. On the 50-task RoboTwin benchmark with 150 episodes per configuration, FlowLap improves success over the asynchronous reference from 88.7% to 91.3% on FlowWAM and, with calibration, from 85.3 % to 92.0 % on LingBot-VA. Relative to matched reduced-work synchronous baselines, it reduces mean transition wait from 896.6 to 39.1 ms on FlowWAM and from 839.9 to 71.8 ms on LingBot-VA, with 0 ms p90 wait for both models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.