RCViT: Encoding Spatial Structure through Alternating Row–Column Attention
Abstract
Positional encodings provide Vision Transformers with information about the spatial arrangement of image patches. Spatial relationships can also be specified through the attention pattern itself: row and column connections represent horizontal and vertical alignment, and alternating these connections across layers composes dependencies along both image axes. We investigate this observation through \method, a Vision Transformer with sparse row–column attention. Deterministic masks combine local and dilated neighbors, alternate the active axis across layers, and assign directional roles to attention heads. Global heads preserve unrestricted interactions, controlling the strength of the spatial inductive bias. The construction adds no trainable parameters and integrates with absolute, relative, conditional, and rotary positional encodings. Classification experiments demonstrate the utility of the attention pattern without positional embeddings and its complementarity with positional encoding in mixed directional–global models. Ablations and resolution and layout evaluations identify the importance of axis alternation and the dependence on head allocation and model configuration. The results support alternating row–column attention as a spatial inductive bias complementary to global attention and positional encoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.