AlignTrack: Aligning Semantics, Geometry, and Actions for Embodied Visual Tracking
Abstract
Embodied visual tracking (EVT) requires agents to jointly understand target se- mantics, reason about spatial relations, and generate actions for continuous target following. While vision-language-action (VLA) models have shown strong poten- tial for EVT, their implicit action prediction leaves semantic, geometric, and action reasoning insufficiently coupled. To address this issue, we propose AlignTrack, a lightweight VLA that explicitly bridges these representations through a semantic- aligned depth encoder and an action-relevant understanding decoder. A two-stage strategy first aligns depth features with pretrained visual-semantic representations and then jointly supervises target-centric spatial understanding and action predic- tion. Extensive experiments on EVT-Benchmark show that AlignTrack with only a 0.6B backbone achieves performance comparable to the 7B TrackVLA, demon- strating the effectiveness of explicit spatial knowledge for lightweight embodied tracking. The code and trained models will be released publicly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.