Adaptive Token–Block Mixing for Efficient Vision Encoders
Abstract
Efficient visual representation learning requires fine-grained token selection and access to long-range context. Local attention restricts direct cross-region interaction, whereas global associative aggregation can limit contextual diversity. We introduce Adaptive Token-Block Mixing (ATBM), a visual mixing operator that combines within-block softmax attention with structured communication over block-specific key-value states. A head-specific Toeplitz mixer propagates these states using learned functions of relative block-index distance, after which token queries retrieve cross-block context. Content-dependent gates fuse the local and cross-block outputs for each token and attention head. This design preserves explicit token interaction within blocks while avoiding a dense global token-to-token attention matrix. On ImageNet-1K, the tiny version ATBM-T achieves 77.2% top-1 accuracy with 6M parameters and 1.2 GFLOPs, improving over DeiT-T by 5.0 percentage points at the same reported model size. At resolution, ATBM-T achieves a inference speedup over DeiT-T while reducing peak GPU memory usage by 82.3%. Evaluations in plain and hierarchical encoders, image generators, together with transfer experiments on COCO and ADE20K, further demonstrate the applicability of the proposed operator.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.