A Simplified DINOv3-Driven Discriminative Framework with Siamese Template Fusion for Visual Tracking
Abstract
Visual object tracking requires representations that separate target from background with high precision while remaining cheap enough for real-time deployment. On UAV platforms, model size and latency constraints make this trade-off especially acute. Lightweight vision transformers achieve the required speed but their limited capacity weakens discriminative power, whereas self-supervised architectures such as DINOv3 provide rich semantic hierarchies that remain underexplored in tracking and, lacking an official tiny variant, cannot be deployed directly on constrained devices. We present DinoSDTrack, which reconciles these two regimes. First, we introduce a simplified, unified Siamese-discriminative framework in which the Siamese template prior is fused into a discriminative tracking head within a one-stream architecture, enabling target-background separation to be learned jointly rather than assembled from independent modules. Second, we show that DINOv3-Base fine-tuned on tracking objectives supplies task-adapted representations, and we formulate a task-adaptive distillation scheme that transfers these representations into a custom DINOv3-Tiny, preserving discriminative accuracy under strict real-time budgets. Experiments on six general tracking benchmarks and six UAV benchmarks show that DinoSDTrack attains state-of-the-art accuracy while maintaining real-time inference speed, indicating that task-adapted self-supervised hierarchies can be effectively compressed for discriminative tracking. Code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.