Spectral Tokenization for Scalable Vision Transformers
Abstract
At , a standard ViT converts RGB patches into tokens; processing a finer image with similar encoder cost requires a different way to construct those tokens. Inspired by the blockwise discrete cosine transform (DCT) used in JPEG compression, we introduce SC-ViT, which tokenizes selected coefficients from a dense grid of image blocks instead of embedding RGB patches directly. It retains a limited spectral bandwidth while sampling the image at : with coefficients per channel and block, the retained tensor has the same number of scalars as a RGB image, and both models form image tokens. Separate luminance and chrominance projections and a finer path for DC coefficients form the token embeddings. After epochs on ImageNet-1K, SC-ViT reaches top 1 accuracy at GMACs, versus at GMACs for ViT-B/16 trained with the same recipe. A separate direct epoch run reaches . A progressive configuration yields three inference budgets from one checkpoint, reaching at GMACs. On KADID-10k, no reference quality prediction reaches Spearman correlation, versus for a spatial ViT-B/16 at . The spectral tokenizer also fits a hierarchical Swin encoder, where SC-Swin reaches on ImageNet-1K after training epochs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.