TRAPE: Compact Axial Position Tapes for Cross-Resolution Vision Transformers
Abstract
Vision transformers trained at one resolution often degrade at larger resolutions or other aspect ratios. We propose TRAPE (Tabulated Relative Axial Positional Encoding), a relative attention bias trained as a continuous function and deployed as compact tables. Along each axis, TRAPE adds a local correction table to a continuous function of log-spaced offsets with a fixed reference; a shared learnable amplitude bounds the sum of both axis responses. The fixed reference makes a smaller grid's table an exact centre slice of a larger one. Fine-tuning only positional parameters at a few square resolutions yields a resolution tape of per-axis tables that height and width slice and blend independently, serving unseen resolutions and aspect ratios with one batched lookup and no positional network. On ImageNet-1k, TRAPE-DeiT-S trained at retains 70.2% top-1 at without fine-tuning, against 36.4% for resized absolute embeddings and 12.1% for 2D RoPE-Mixed. As a zero-initialized branch, TRAPE raises five frozen public ViTs by 18–56 points at with under 0.1 ImageNet epoch of position-only adaptation. On one A800 GPU, assembling tables for a new shape takes about 0.1 ms. Fused attention avoids materializing the bias matrix, enabling over the batch size of dense attention at . Hierarchical TRAPE backbones with axial attention improve ADE20K segmentation by 3.0–3.1 mIoU over Swin with fewer FLOPs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.