acceptodds
Under review as a conference paper at ICLR 2027

GeoVLT: Camera Geometry Conditioning for UAV Vision-Language Tracking

Abstract

In UAV-based urban monitoring, continuous target tracking often faces changes in altitude, gimbal pitch angle and focal length, resulting in different camera geometry settings that affect target imaging and scene coverage. Most existing UAV vision-language trackers rely on learned features to implicitly handle the resulting imaging changes, potentially reducing tracking reliability across geometry settings. Camera geometry provides physical cues about these imaging changes, enabling more reliable target modeling. Therefore, we introduce GeoVLT, a vision-language tracking framework that uses these cues throughout the tracking process. Specifically, Geometry-Conditioned Representation first adapts the target representation within a single frame. Geometry-Adaptive Temporal Propagation then maintains historical target information through multi-frame propagation. Geometry-Conditioned Decoding finally combines the single-frame representation with multi-frame historical information for coordinate prediction. To evaluate the use of camera geometry in tracking, we construct GeoUAV, a dataset that provides vision, language, and camera geometry information. For public UAV benchmarks without camera geometry, we train a lightweight estimator on GeoUAV to predict camera geometry from first-frame features. GeoVLT achieves state-of-the-art performance among the evaluated trackers, reaching 69.69% AUC on GeoUAV100 compared with 67.31% for ATCTrack, and 64.11% AUC on UAV20L-NLP compared with 61.66% for UITrack. These results show the effectiveness of explicitly using camera geometry in vision-language tracking. Our code and checkpoints are released at https://anonymous.4open.science/r/GeoVLT.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.