acceptodds
Under review as a conference paper at ICLR 2027

Spectral Tokenization for Scalable Vision Transformers

Abstract

At , a standard ViT converts RGB patches into tokens; processing a finer image with similar encoder cost requires a different way to construct those tokens. Inspired by the blockwise discrete cosine transform (DCT) used in JPEG compression, we introduce SC-ViT, which tokenizes selected coefficients from a dense grid of image blocks instead of embedding RGB patches directly. It retains a limited spectral bandwidth while sampling the image at : with coefficients per channel and block, the retained tensor has the same number of scalars as a RGB image, and both models form image tokens. Separate luminance and chrominance projections and a finer path for DC coefficients form the token embeddings. After epochs on ImageNet-1K, SC-ViT reaches top 1 accuracy at GMACs, versus at GMACs for ViT-B/16 trained with the same recipe. A separate direct epoch run reaches . A progressive configuration yields three inference budgets from one checkpoint, reaching at GMACs. On KADID-10k, no reference quality prediction reaches Spearman correlation, versus for a spatial ViT-B/16 at . The spectral tokenizer also fits a hierarchical Swin encoder, where SC-Swin reaches on ImageNet-1K after training epochs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.