FLOYD VISION TRANSFORMER: RELATION-CENTRIC VISUAL BACKBONES WITH PIVOTAL ATTENTION
Abstract
As vision models are increasingly applied to complex real-world tasks, visual understanding must extend beyond recognizing individual objects to modeling their relationships, interactions, and scene dynamics. Mainstream visual backbones typically operate on patches, leaving relations between patches implicit in the resulting token representations. Pivotal Attention (PA) makes relations explicit through ordered-pair features and their updates, but its all-pairs computation is costly for images. We introduce Row Column Pivotal Attention (RCPA), which exploits the organized image grid to jointly select and aggregate row and column features. This structure preserves global two-hop coverage through sparse connections, as characterized by the rook's graph, while reducing attention interactions from PA's to for patches on a square grid. Building on RCPA, we develop Vision Transformer Floyd (ViF), a general-purpose backbone that brings PA-inspired relation modeling to visual recognition and dense prediction. On ClockBench relative-angle prediction, PA and RCPA Small reduce IID angular error by 83.6% and 67.3% relative to ViT Small. On real images, ViF Base improves VG150 SGDet R@100 by 4.6% relative to ViT Base and transfers effectively to ImageNet-1K, ADE20K, and COCO. These results establish a practical route from explicit relation modeling to visual backbone design and motivate its extension to richer relational and multimodal tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.