acceptodds
Under review as a conference paper at ICLR 2027

Where to Interact and When to Skip: Efficient Transformer for Underwater Visual Tracking

Abstract

Underwater visual tracking is challenged by degraded imaging conditions such as light attenuation and scattering, as well as by limited computational resources on underwater platforms that restrict real-time deployment. Existing Transformer trackers typically perform dense template–search interactions and execute all Transformer layers for every frame, resulting in considerable redundant computation. To address these issues, we propose WTrack, an efficient Transformer-based underwater tracker that allocates computation adaptively in both token interaction and network depth, determining where token interactions should occur and when redundant Transformer blocks can be skipped. Specifically, WTrack contains two key components: (i) Stage-aware Attention Interaction (SAI), which reorganizes template-search interactions across Transformer stages. In shallow stages, SAI removes search-to-template interactions to prevent ambiguous information propagation caused by dense and visually similar underwater targets. In deeper stages, SAI eliminates redundant intra-frame interactions and focuses on bidirectional template-search correspondence to improve target matching and localization. (ii) Adaptive Layer Execution (ALE), which dynamically adjusts Transformer depth according to the tracking state. Since underwater sequences exhibit varying target visibility, motion patterns, and image quality, ALE estimates scene complexity and skips redundant deep-layer computation in easy cases where additional layers provide limited benefits. Experiments on UTB180, UOT100, and VMAT demonstrate that WTrack achieves competitive tracking accuracy and real-time efficiency, reaching 210.52 FPS on GPU while outperforming existing real-time trackers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.