acceptodds
Under review as a conference paper at ICLR 2027

Ambiguous Token Refinement for Target Representation in Transformer Tracking

Abstract

Transformer-based visual trackers have recently achieved strong performance through effective template–search interaction. To mitigate background interference in template modeling, existing methods commonly introduce token-type embeddings to separate foreground and background tokens, treating all tokens inside the target bounding box as homogeneous foreground. However, such a coarse binary strategy relies solely on geometric overlap and ignores the semantic content within the box, causing background-like tokens inside the target region to confuse the template representation and degrade tracking performance. In this paper, we revisit template modeling in Transformer-based tracking and propose a general ambiguous token refinement method that refines token types inside the target bounding box. It is model-agnostic, and can be readily applied to various tracking tasks, including RGB, RGB+Thermal, RGB+Depth, and RGB+Event tracking. Specifically, unlike the previous coarse binary foreground/background template tokens strategy, our approach automatically assigns an additional ambiguous type to template tokens that are semantically similar to the surrounding background, thereby reducing confusion in template representations. Despite its simplicity, extensive experiments on visible and multi-modal benchmarks demonstrate consistent improvements over strong baselines and outperform many state-of-the-art trackers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.