acceptodds
Under review as a conference paper at ICLR 2027

TRAPE: Compact Axial Position Tapes for Cross-Resolution Vision Transformers

Abstract

Vision transformers trained at one resolution often degrade at larger resolutions or other aspect ratios. We propose TRAPE (Tabulated Relative Axial Positional Encoding), a relative attention bias trained as a continuous function and deployed as compact tables. Along each axis, TRAPE adds a local correction table to a continuous function of log-spaced offsets with a fixed reference; a shared learnable amplitude bounds the sum of both axis responses. The fixed reference makes a smaller grid's table an exact centre slice of a larger one. Fine-tuning only positional parameters at a few square resolutions yields a resolution tape of per-axis tables that height and width slice and blend independently, serving unseen resolutions and aspect ratios with one batched lookup and no positional network. On ImageNet-1k, TRAPE-DeiT-S trained at retains 70.2% top-1 at without fine-tuning, against 36.4% for resized absolute embeddings and 12.1% for 2D RoPE-Mixed. As a zero-initialized branch, TRAPE raises five frozen public ViTs by 18–56 points at with under 0.1 ImageNet epoch of position-only adaptation. On one A800 GPU, assembling tables for a new shape takes about 0.1 ms. Fused attention avoids materializing the bias matrix, enabling over the batch size of dense attention at . Hierarchical TRAPE backbones with axial attention improve ADE20K segmentation by 3.0–3.1 mIoU over Swin with fewer FLOPs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.