Are Vision-Language Agents Latency-Aware? A Systematic Evaluation Under Action Latency
Abstract
Foundation-model agents, such as vision-language-action (VLA) policies, suffer from inference latency: the environment keeps changing while the policy computes. Faster inference reduces this delay but does not eliminate it, and training with a delay chosen in advance may not match the target deployment. Current benchmarks hide the problem by pausing the environment during inference; evaluated in real time on an RTX 3090 GPU, VLA policies on 12 environments retain a median of only 6.6% of their zero-latency return. We study how to estimate the residual latency of a deployment, adapt a policy to it, and determine when the adaptation transfers. We introduce a **latency profiling and replay framework** that fits the distribution and temporal dependence of measured action latencies and replays new delay sequences in simulation. (1) Across five GPU deployments, a Temporal profile reproduces hardware-timed return with an average gap of 10.8, versus 87.4 for a fixed mean delay. (2) Training under the target profile improves hardware-timed return in all 36 task–architecture pairs, by 16% to 86% of zero-latency return, while transfer across delays is asymmetric and task-dependent. (3) An explicit latency value can be essential when training across delays, whereas the benefit of visual history depends on the task. Residual deployment latency can be modeled and incorporated into policy training, although adaptation does not generalize uniformly across latency regimes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.