acceptodds
Under review as a conference paper at ICLR 2027

Keyword Enhanced Generative Priors for Vision Language Tracking in Underwater Conditions

Abstract

Vision Language Tracking (VLT) localizes target objects in videos by jointly exploiting visual and textual cues, ensuring stronger reliability than traditional tracking methods. However, existing VLT methods still suffer from severe performance degradation in underwater scenarios because the strong noise and color distortion resemble the target and its surroundings closely. Additionally, the lack of corresponding training data prevents this limitation from being addressed through a pure data-driven approach. In this paper, we address this limitation through a more efficient method, transferring the strong prior of generative models learned from the pretraining on large-scale natural images to underwater scenarios. Specifically, we incorporate a keyword enhancement strategy. It extracts intermediate multi-modal features from the generation process to assist tracking. During this, it extracted the keywords related to the target and modifies cross-attention maps by increasing the corresponding attention weights to steer the features to place greater emphasis on the tracking target. With this in mind, we propose KeyTrack, a generative-model-based VLT framework, which uses strong priors from generative models and enables target-focused representation learning while suppressing irrelevant environmental noise, thereby achieving more accurate tracking. Extensive experiments demonstrate that KeyTrack achieves state-of-the-art performance on underwater tracking benchmarks and maintains competitive results on general datasets, validating its robustness and generalization capability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.