Every Sparkle Gets a Token: Entropy-Guided Foveated Tokenization for Fine-Grained Visual Recognition
Abstract
Uniform image tokenization allocates the same number, locations, and scales of tokens to every image, even when discriminative evidence is concentrated within a small region. Existing token-reduction methods commonly identify redundancy only after a dense patch sequence has entered the backbone. We introduce SpiralFovea, a transparent, parameter-free selector that constructs a sparse visual sequence directly from image content. SpiralFovea uses local luminance entropy to identify spatially distributed anchors and extracts multi-scale rings of crops around them, allocating more tokens to locally informative regions while retaining coarser contextual coverage. These rings also define a deterministic centre-to-periphery ordering suitable for directional state-space models such as Mamba, while transformers receive explicit spatial descriptors for the resulting non-grid tokens. At 224 × 224 resolution, SpiralFovea uses at most 78 spatial tokens instead of the 196-token ViT-S/16 grid. Across four recognition tasks, it improves the reported mean accuracy of frozen ViT/DINO and ResNet–Mamba pipelines while increasing measured end-to-end throughput by 17.8–28.9%. On standard 200-way CUB at a matched sparse budget, SpiralFovea achieves 85.16% top-1 accuracy—within 0.45 percentage points of the dense model and 0.99 points above ToMe. At 518 × 518 resolution, it reduces a 1,369-token DINOv2 grid to approximately 120–200 tokens, revealing substantial computational savings but task-dependent accuracy trade-offs. Foreground-localization and inverted-entropy analyses further identify the method’s operating boundary: entropy is effective when discriminative structure is locally textured and spatially coherent, but becomes unreliable when high-entropy background clutter dominates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.