Texture as Signal: Spectral Adaptive Tokenization for Vision Transformers
Abstract
Vision Transformers (ViTs) struggle with quadratic computational costs, a bottleneck often reduced by adaptive tokenization. Current methods predominantly rely on spatial sparsity (using edge detection or region homogeneity) to reduce token counts. However, this spatial tokenization fails in texture-dominated regimes (such as high-resolution aerial and microscopic imagery) where high-frequency gradients are dense. In these environments, edge-driven models either over-fragment the image or aggressively merge critical details, destroying both compression and performance. To resolve this, we introduce texture-aware adaptive tokenization. Rather than relying on spatial discontinuities, we formulate token saliency through spectral statistics to capture local frequency content and redundancy. We propose a general framework that dynamically partitions tokens based on texture complexity, requiring no modifications to the underlying ViT architecture. Supported by a theoretical analysis of spectral density failure modes, this approach is extensively evaluated across SAM, Swin, and ViT architectures on diverse, texture-rich datasets. The proposed texture-driven framework achieves up to a 41.5 speedup of training wall-clock time for SAM on wild blueberry aerial imagery, while simultaneously improving convergence dice scores (from 63.1% to 71.2%). Furthermore, under severe image perturbations, the texture partitioning mechanism maintains greater than 91% Structural Intersection over Union (SIoU) under heavy Gaussian noise and greater than 84% under severe Gaussian blur, preventing the catastrophic structural collapse observed in conventional edge-based tokenization (as low as 50.90%). Ultimately, these findings establish that spectral properties (rather than spatial boundaries) are the governors of token saliency in texture-dominated regimes, enabling the efficient deployment of heavy foundation models in complex visual domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.