DART: Differentiable Adaptive Region Tokenization for Unified Visual Backbones
Abstract
A 29M-parameter Vim-Small equipped with our tokenizer surpasses a 98M-parameter Vim-Base on ImageNet (82.2% vs. 81.9%) at half the computational cost (10.9G vs. 19.9G FLOPs). We achieve this through a simple allocation principle: each token should cover an equal share of learned importance rather than an equal share of area. We realize this principle through DART, a closed-form construction that inverts a piecewise-linear CDF at uniform quantile points. The resulting tokenizer is differentiable, always outputs a fixed-length regular tensor, and recovers the standard fixed grid exactly when the importance map is uniform. Under the standard 196-token budget, DART consistently improves accuracy by +0.8% to +1.6% across five backbone configurations spanning Transformers and state space models. The consistency of improvements across architectures with fundamentally different sequence-processing mechanisms suggests the equal-importance principle captures a general property of visual information rather than exploiting architecture-specific biases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.