acceptodds
Under review as a conference paper at ICLR 2027

DEFLECT: Temporal Counterfactual Preference Learning for Delay-Robust Asynchronous VLAs

Abstract

Vision-Language-Action (VLA) policies increasingly rely on asynchronous inference to hide model latency behind ongoing robot motion. During inference, the robot continues executing its current action chunk while the next chunk is generated from the observation available at inference start. By the time the new chunk is executed, the robot or scene may have changed, so an action that fits the prediction-time state can be misaligned with the execution-time state. Training-time methods modify the policy input or supervised target based on the expected execution delay, but do not directly supervise how actions should change as the scene evolves between prediction and execution. Runtime repair methods improve chunk continuity or modify actions during deployment, but can still leave actions aligned with the prediction-time state rather than the execution-time state. Our key insight is that execution-time observations are unavailable when the online prediction is made but are available in recorded trajectories during offline training, making the prediction–execution mismatch observable during training. We propose DEFLECT, an offline post-training framework that turns this temporal contrast into preference supervision for a policy acting from stale inputs. Across Kinetix, LIBERO, and three real-robot tasks, DEFLECT improves delay robustness across the evaluated VLA policies and settings. At the longest supported delay, it improves success by 4.2 percentage points over VLASH on Kinetix and by 22.7 points over naive asynchronous execution with GR00T N1.7 on LIBERO.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.