acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Geometric Perception Kernels

Abstract

Modern vision architectures typically construct visual representations using predefined spatial structures, such as fixed patch grids or convolutional receptive fields. We introduce , a vision architecture that learns a compact set of content-dependent anisotropic Gaussian kernels. Each kernel adapts its location, scale, and orientation to the input image and directly aggregates a feature map into a visual token, allowing them to be processed by a standard Transformer architecture. Across image-classification benchmarks, AdaGPK consistently outperforms fixed-patch Vision Transformer baselines and is competitive with or exceeds the evaluated convolutional baselines on ImageNet-scale classification. Compared with unconstrained adaptive pooling, Gaussian parameterization preserves classification performance while producing more spatially structured representations and substantially stronger cross-dataset zero-shot localization. Analysis of the learned representation shows that increasing the kernel budget produces progressively smaller, less overlapping spatial supports and more differentiated token features, while the learned geometry remains stable across augmentations and training seeds. Together, these results suggest that learned spatial decomposition provides an effective way to construct compact visual representations with explicit spatial support.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.