What-Where Transformer: Concurrent Representation and Localization in a Slot-Centric Visual Backbone
Abstract
Image understanding tasks involve identifying what is present and where it appears. However, classification-oriented visual backbones primarily optimize semantic representations of what, while spatial information about where often remains implicit or entangled with semantic features. We hypothesize that explicit separation of these what and where representations can induce localization ability within backbones. To examine this, we introduce the What-Where Transformer (WWT), a ViT-style attentive backbone that incorporates what-where separation as an architectural inductive bias. Our method introduces two key designs: (1) it treats tokens as representations of what and attention maps as representations of where, and processes them in concurrent feed-forward modules via a multistream, slot-based architecture; (2) it reuses both the final-layer tokens and attention maps for downstream tasks, and directly exposes them to gradients derived from tasks, facilitating explicit learning of localization. Trained on ImageNet with single-label classification and an auxiliary reconstruction objective, WWT discovers multiple objects directly from its raw attention maps without token clustering or other postprocessing, at a modest cost to classification accuracy. Furthermore, WWT achieves superior performance compared to ViT-based methods on zeroshot object discovery and weakly supervised semantic segmentation, and it is transferable to various localization setups with minimal modifications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.