acceptodds
Under review as a conference paper at ICLR 2027

The Coupling Form of Positional Encoding Determines Multi-Teacher Generalization in Feature Upsampling

Abstract

Cross-attention feature upsamplers increase the spatial resolution of vision foundation model (VFM) features. Recent upsamplers such as NAF build queries and keys from two sources: image content encoded by a lightweight image encoder, and pixel positions supplied by a rotary position embedding (RoPE). Training such an upsampler with multiple VFMs as teachers is a natural way to improve its transfer to unseen VFMs, yet it has been reported to yield no clear gains, and the cause is not understood. We show that the cause lies in how position is coupled with content. RoPE rotates queries and keys, so position multiplies content. Different teachers require different shapes of the spatial kernel, but under RoPE only the shared image encoder can produce this difference, so the gradients of the teachers conflict within the first 100 training steps. This conflict is resolved not by optimization that modifies the gradients, but by replacing RoPE with a relative position bias. This bias adds position to the attention score as a separate term and learns the kernel shape instead of the image encoder. By building every scheme on NAF and swapping only the positional encoding, on Cityscapes semantic segmentation, adding a second teacher VFM causes a 0.83 mIoU drop for RoPE compared to its best single-teacher baseline, whereas the schemes that add position to the attention score gain up to 0.57 mIoU. An upsampler that learns from multiple VFMs should add position to the attention score instead of multiplying it with content.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.