FalconMOT: Self-Calibrating Evidential Association and Visual-Kinematic Recovery for Aerial Multi-Object Tracking
Abstract
Multi-object tracking (MOT) aims to maintain continuous target trajectories across video frames. In unmanned aerial vehicle (UAV) recorded videos, rapid platform motion, continuous viewpoint variations, and small object scales pose severe challenges, frequently causing unreliable affinity estimation and detector dropouts. Existing methods typically rely on fixed linear weighting between motion and appearance, which fails to adapt when visual cues fluctuate, or employ unconstrained feature-map querying that triggers false recoveries from visually similar distractors. In this work, we propose FalconMOT, a robust aerial tracking framework driven by two training-free components: Self-calibrating Probabilistic Evidence Association (SPEAR) and Query Appearance-Motion (QAM) target recovery. Specifically, SPEAR dynamically balances motion and appearance by evaluating the directional consistency of historical track embeddings; when visual features fluctuate, its concentration parameter automatically decreases, down-weighting appearance affinity and relying on spatial motion for robust association. The QAM module addresses missed detections by querying dense backbone feature maps around predicted track locations, filtering out ambiguous background distractors to ensure reliable recovery. Extensive experiments on VisDrone2019 and UAVDT demonstrate that FalconMOT outperforms current state-of-the-art methods, achieving 69.1% IDF1 and 55.6% MOTA while operating at 25.0 FPS on an embedded Jetson Orin NX in a plug-and-play, training-free manner.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.