acceptodds
Under review as a conference paper at ICLR 2027

Spatial-Frequency Collaborative Mamba Tracker for Visual Object Tracking

Abstract

Visual object tracking has achieved substantial progress, yet robust tracking in complex dynamic scenes remains chal lenging because spatial-domain appearance representations are vulnerable to fast motion, illumination variations, oc clusions, and background distractors. Existing Mamba-based trackers improve long-range dependency modeling and com putational efficiency, but mainly focus on spatial-temporal representations, leaving the temporal evolution of frequency domain motion information insufficiently explored. To ad dress this limitation, we propose a spatial-frequency collab orative tracking framework based on Mamba, termed SFMT, which jointly models spatial-temporal appearance representa tions and frequency-domain motion cues. First, a Frequency Domain Awareness (FDA) module employs Mamba to model the temporal evolution of phase spectra and uses the re sulting motion cues to guide amplitude filtering. Second, a Spatial-Frequency Fusion (SFF) module performs global sequence-level and local point-level interaction to effectively integrate complementary information across the two domains. Finally, a pyramid multi-scale decoder enhances high- and low-frequency representations and dynamically balances their contributions through gated adaptive fusion for accurate tar get localization. Extensive experiments on six challenging benchmarks demonstrate the effectiveness of SFMT, achiev ing competitive tracking performance, including an AUC of 74.1% on LaSOT and an AO of 77.2% on GOT-10K.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.