DASTrack: Exploiting DINOv3 Features, Adaptive Modulation, and Spatial Refinement for Visual Tracking
Abstract
Visual object tracking has benefited from adopting pre-trained Vision Transformers as backbones. While the recently introduced DINOv3 offers rich intermediate-layer representations, it has not been specifically adapted for tracking, and existing backbone adaptation strategies are insufficient to fully exploit its potential. Moreover, existing feature modulation strategies rely on template-search interaction and coordinate-based trajectory encoding, lacking explicit channel-wise conditioning on target appearance and trajectory modulation that accounts for spatial response distributions. To address these limitations, we propose DASTrack, a framework that exploits DINOv3 features for tracking through dedicated backbone adaptation, adaptive modulation, and spatial refinement. For backbone adaptation, we propose dense LoRA, which jointly encodes template and search images and aggregates multi-layer features, exploiting reliable intermediate representations of the backbone. To explicitly condition features on the target, we introduce template-adaptive modulation and trajectory-guided modulation to enhance search features based on target appearance and historical response distributions, respectively. We further design a spatial refinement head that iteratively refines search features through Gaussian-gated attention to achieve precise localization. We evaluate DASTrack on seven benchmarks, achieving new state-of-the-art results of 81.4% AO on GOT-10k and 72.6% AUC on UAV123, with competitive results on other datasets. Code and results will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.