VisionAlphabet: Compressing Spatial Context with Learned Laplace Poles
Abstract
VisionAlphabet introduces a distinct approach to compact visual backbones, whereby spatial context is represented and propagated through learned Laplace poles. Each pole provides a stable complex-valued state that carries spatial context: damping controls how far contributions persist across space, while phase modulates how contributions from different locations combine. Four directional scans propagate spatial evidence across the image; the terminal scan's gain-normalized pole energies are spatially averaged to form the image-level classifier descriptor. Under a shared 100-epoch ImageNet-1K protocol without distillation, the 3.25M- and 5.08M-parameter models attain 72.35% and 75.02% Top-1 accuracy, respectively, while using 10–12% fewer parameters than their corresponding compact baselines. The larger variant also outperforms the compared baselines with 10% of the ImageNet-1K training data and, when pretrained on full ImageNet-1K, transfers to COCO detection and instance segmentation. Controlled interventions on a frozen checkpoint show that recognition changes when pole-state transport is disrupted or phase is removed, supporting the interpretation of pole states as functional spatial memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.