acceptodds
Under review as a conference paper at ICLR 2027

Exact Native Vertices for Continuous Adaptive Tokenization

Abstract

Choosing where to use fine patches is central to adaptive image tokenization, but each choice changes the length and structure of the token sequence. We introduce Differentiable Virtual-Slot Partitioning (DVSP), which places coarse and fine representations in a shared, fixed-length computation graph. Each image region is represented by four virtual slots whose content, position, and attention mass vary together with a learned resolution coordinate. At the binary endpoints, these slots reproduce either one coarse token or four fine tokens. We prove that every resulting binary partition exactly matches its native variable-length execution under compatible deterministic transformer operations. This allows resolution to be learned through task gradients, with projection evaluated separately from exact execution. On ImageNet, DVSP reduces visual-token counts by 30.4% on ViT-Small with a Top-1 accuracy decrease of 0.53 percentage points. Controlled spatial comparisons show that learned partitions outperform entropy selection at identical per-image token counts and fixed task weights. DVSP makes adaptive input resolution learnable within a continuous representation whose discrete states retain their native meaning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.