acceptodds
Under review as a conference paper at ICLR 2027

AlignTrack: Aligning Semantics, Geometry, and Actions for Embodied Visual Tracking

Abstract

Embodied visual tracking (EVT) requires agents to jointly understand target se- mantics, reason about spatial relations, and generate actions for continuous target following. While vision-language-action (VLA) models have shown strong poten- tial for EVT, their implicit action prediction leaves semantic, geometric, and action reasoning insufficiently coupled. To address this issue, we propose AlignTrack, a lightweight VLA that explicitly bridges these representations through a semantic- aligned depth encoder and an action-relevant understanding decoder. A two-stage strategy first aligns depth features with pretrained visual-semantic representations and then jointly supervises target-centric spatial understanding and action predic- tion. Extensive experiments on EVT-Benchmark show that AlignTrack with only a 0.6B backbone achieves performance comparable to the 7B TrackVLA, demon- strating the effectiveness of explicit spatial knowledge for lightweight embodied tracking. The code and trained models will be released publicly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.