acceptodds
Under review as a conference paper at ICLR 2027

DART: Differentiable Adaptive Region Tokenization for Unified Visual Backbones

Abstract

A 29M-parameter Vim-Small equipped with our tokenizer surpasses a 98M-parameter Vim-Base on ImageNet (82.2% vs. 81.9%) at half the computational cost (10.9G vs. 19.9G FLOPs). We achieve this through a simple allocation principle: each token should cover an equal share of learned importance rather than an equal share of area. We realize this principle through DART, a closed-form construction that inverts a piecewise-linear CDF at uniform quantile points. The resulting tokenizer is differentiable, always outputs a fixed-length regular tensor, and recovers the standard fixed grid exactly when the importance map is uniform. Under the standard 196-token budget, DART consistently improves accuracy by +0.8% to +1.6% across five backbone configurations spanning Transformers and state space models. The consistency of improvements across architectures with fundamentally different sequence-processing mechanisms suggests the equal-importance principle captures a general property of visual information rather than exploiting architecture-specific biases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.