Keyword Enhanced Generative Priors for Vision Language Tracking in Underwater Conditions
Abstract
Vision Language Tracking (VLT) localizes target objects in videos by jointly exploiting visual and textual cues, ensuring stronger reliability than traditional tracking methods. However, existing VLT methods still suffer from severe performance degradation in underwater scenarios because the strong noise and color distortion resemble the target and its surroundings closely. Additionally, the lack of corresponding training data prevents this limitation from being addressed through a pure data-driven approach. In this paper, we address this limitation through a more efficient method, transferring the strong prior of generative models learned from the pretraining on large-scale natural images to underwater scenarios. Specifically, we incorporate a keyword enhancement strategy. It extracts intermediate multi-modal features from the generation process to assist tracking. During this, it extracted the keywords related to the target and modifies cross-attention maps by increasing the corresponding attention weights to steer the features to place greater emphasis on the tracking target. With this in mind, we propose KeyTrack, a generative-model-based VLT framework, which uses strong priors from generative models and enables target-focused representation learning while suppressing irrelevant environmental noise, thereby achieving more accurate tracking. Extensive experiments demonstrate that KeyTrack achieves state-of-the-art performance on underwater tracking benchmarks and maintains competitive results on general datasets, validating its robustness and generalization capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.