Stream-in-Stream: Teaching Vision-Language-Action Models when Their Actions Expire
Abstract
Vision-language-action (VLA) models are mostly trained and evaluated in scenes that hold still while they compute. We study the stream-in-stream setting, in which the world keeps moving during inference, and make action timing measurable and learnable. PhysClock adds granular, settling, rolling, and air-coupled task fami- lies to simulated manipulation and advances the world while a policy computes; PhysClock-Syn labels scripted demonstrations with replay-measured timing: how much delay each decisive action tolerates, and how early and how late each action chunk may start. Fine-tuned on these demonstrations and run synchronously, π0.5 still gains 14 percentage points when the world pauses during its inference. We train a small head on the timing labels to predict the earliest and latest start of each chunk, and a scheduler waits for, keeps, or discards chunks accordingly. In a confirmatory evaluation with contrasts fixed in advance, the head raises success on the three families other than a granular pour by 3.6 and 5.5 points on two poli- cies, mostly in episodes with little timing tolerance, also when attached to frozen policies; a head that predicts success probabilities from the same labels is equiv- alent to it within a pre-specified margin. On a real arm, heads trained on event times read from demonstration videos raise grasps of a falling sheet from 58 to 83 of 200 paired trials in a pre-registered batch, after a smaller one was inconclusive, and catches of a rolling ball from 41 to 60 of 120. A head trained in simulation does not transfer to the real arm, and on a task family held out from training the head brings no measurable gain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.