acceptodds
Under review as a conference paper at ICLR 2027

On the move: Latency-aware Embodied Visual Tracking

Abstract

Embodied visual tracking (EVT) is inherently latency-sensitive. While a robot computes its next action, the target keeps moving, so observations grow stale before actions take effect. Existing benchmarks, however, advance the environment only after the policy returns its action, charging nothing for decision time. We introduce RT-EVT, a latency-aware tracking benchmark that evaluates policies under configurable action latency. Accounting for deployment latency substantially changes tracking performance. For example, LightNav-0's tracking success rate drops from a SOTA 89.8% to 22.9% under 200 ms latency, even worse than a 5× smaller model. We propose TrackRT, a compact 0.9B vision-language navigation model designed for real-time tracking. We further introduce latency-aware training, implemented through action-effect-state supervision. The policy is trained against expert trajectories from the states in which its actions take effect. On the Jetson Orin NX, TrackRT achieves a mean inference latency of 145.0 ms versus 1201.0 ms for LightNav-0, with corresponding RT-EVT STT success rates of 65.3% and 0.5% respectively. Controlled ablations across different latency settings further confirm the benefit of latency-aware training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.