Learning to See through Mnemonic Recall
Abstract
Modern visual backbones largely follow an operator-centric view of representation learning, where prior experience shapes the transformations applied to inputs. Yet experience can also contribute as persistent visual content that is selectively recalled to inform current perception. Inspired by memory-based reinstatement in human perception, we formulate visual representation learning as the joint contribution of dynamic computation and conditional recall. This work introduces ReMeViT, a family of visual Transformer backbones equipped with visual pattern memory that retains recurring local patterns as persistent and addressable states. Current visual cues can selectively evoke relevant patterns, while context-conditioned gating adapts their contribution to the current observation. The separation of stored content from contextual use also allows a shared source of visual experience across different backbone architectures. We instantiate ReMeViT on both DeiT and Swin, which yield consistent performance gains across classification and dense prediction tasks. Further results support visual pattern memory as a source of reusable perceptual priors that benefit new tasks, strengthen recognition under occlusion, and enhance representations learned without supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.