acceptodds
Under review as a conference paper at ICLR 2027

VLA-HWBench: Where the Time Goes When VLA Models Run on Edge Hardware

Abstract

Vision–language–action (VLA) models increasingly run on edge devices aboard robots, where inference latency constrains closed-loop robot control. However, it remains unclear how to design VLA models, their inference software and edge hardware so that robots can run VLA inference in real time. Most VLA models report only task success, and prior studies cover few of the ways recent VLAs generate actions, split a call into at most three phases and do not compare inference optimizations on the same models. Therefore, we introduce **VLA-HWBench**, a benchmark that measures 18 open VLA models on their released software and on alternative backends, on an NVIDIA Jetson AGX Thor and an AMD Ryzen AI Max+ PRO 395, with a seven-stage breakdown of each call. Our study reveals three insights. First, where the time goes depends on how a model generates actions, not on its size. Action generation is the largest stage in 14 of the 19 configurations on the Jetson, latency spans on the Jetson and on the AMD platform, and parameter count explains at most 15% of its variance. Models that generate long action chunks in a few flow-matching steps keep up with control, whereas token-by-token decoding does not. Second, kernel fusion with 4-bit weights helps most, running and GR00T N1.7 2.6–2.9 faster than PyTorch with torch.compile, but it exists for only these two models and only on the Jetson. Optimization also moves the time to the stages it shortens least, as CPU preprocessing grows from 12% to 39% of the GR00T N1.7 call. Third, after optimization, memory bandwidth limits only the stages that process few tokens on the Jetson, such as the denoising steps, whereas software limits all other stages, from unfused kernels to serial preprocessing on the host CPU. Overall, low-latency edge VLAs should use few action-generation iterations, inference software should extend kernel fusion and low-bit weights to more models and platforms, and edge hardware needs high memory bandwidth and a fast host CPU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.