acceptodds
Under review as a conference paper at ICLR 2027

RCViT: Encoding Spatial Structure through Alternating Row–Column Attention

Abstract

Positional encodings provide Vision Transformers with information about the spatial arrangement of image patches. Spatial relationships can also be specified through the attention pattern itself: row and column connections represent horizontal and vertical alignment, and alternating these connections across layers composes dependencies along both image axes. We investigate this observation through \method, a Vision Transformer with sparse row–column attention. Deterministic masks combine local and dilated neighbors, alternate the active axis across layers, and assign directional roles to attention heads. Global heads preserve unrestricted interactions, controlling the strength of the spatial inductive bias. The construction adds no trainable parameters and integrates with absolute, relative, conditional, and rotary positional encodings. Classification experiments demonstrate the utility of the attention pattern without positional embeddings and its complementarity with positional encoding in mixed directional–global models. Ablations and resolution and layout evaluations identify the importance of axis alternation and the dependence on head allocation and model configuration. The results support alternating row–column attention as a spatial inductive bias complementary to global attention and positional encoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.