On the move: Latency-aware Embodied Visual Tracking
Abstract
Embodied visual tracking (EVT) is inherently latency-sensitive. While a robot computes its next action, the target keeps moving, so observations grow stale before actions take effect. Existing benchmarks, however, advance the environment only after the policy returns its action, charging nothing for decision time. We introduce RT-EVT, a latency-aware tracking benchmark that evaluates policies under configurable action latency. Accounting for deployment latency substantially changes tracking performance. For example, LightNav-0's tracking success rate drops from a SOTA 89.8% to 22.9% under 200 ms latency, even worse than a 5× smaller model. We propose TrackRT, a compact 0.9B vision-language navigation model designed for real-time tracking. We further introduce latency-aware training, implemented through action-effect-state supervision. The policy is trained against expert trajectories from the states in which its actions take effect. On the Jetson Orin NX, TrackRT achieves a mean inference latency of 145.0 ms versus 1201.0 ms for LightNav-0, with corresponding RT-EVT STT success rates of 65.3% and 0.5% respectively. Controlled ablations across different latency settings further confirm the benefit of latency-aware training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.