The Coupling Form of Positional Encoding Determines Multi-Teacher Generalization in Feature Upsampling
Abstract
Cross-attention feature upsamplers increase the spatial resolution of vision foundation model (VFM) features. Recent upsamplers such as NAF build queries and keys from two sources: image content encoded by a lightweight image encoder, and pixel positions supplied by a rotary position embedding (RoPE). Training such an upsampler with multiple VFMs as teachers is a natural way to improve its transfer to unseen VFMs, yet it has been reported to yield no clear gains, and the cause is not understood. We show that the cause lies in how position is coupled with content. RoPE rotates queries and keys, so position multiplies content. Different teachers require different shapes of the spatial kernel, but under RoPE only the shared image encoder can produce this difference, so the gradients of the teachers conflict within the first 100 training steps. This conflict is resolved not by optimization that modifies the gradients, but by replacing RoPE with a relative position bias. This bias adds position to the attention score as a separate term and learns the kernel shape instead of the image encoder. By building every scheme on NAF and swapping only the positional encoding, on Cityscapes semantic segmentation, adding a second teacher VFM causes a 0.83 mIoU drop for RoPE compared to its best single-teacher baseline, whereas the schemes that add position to the attention score gain up to 0.57 mIoU. An upsampler that learns from multiple VFMs should add position to the attention score instead of multiplying it with content.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.