Decoding Register Tokens Preserves Global Scene Details
Abstract
Pretrained vision transformers learn structured image representations spanning spatial patch tokens and non-spatial tokens. Recent analyses suggest that non-spatial register tokens encode global scene properties such as illumination, blur, and style. Yet patch-based image decoders discard these outputs, leaving the decoder to recover image appearance from patches alone. We investigate whether retaining register features improves decoding when the complete spatial patch representation is already available. We study this question using representation autoencoders (RAEs) with frozen register-equipped encoders. Our decoder processes four register tokens together with all patch tokens through joint self-attention. Appearance probes, paired reconstruction-error analysis, and fixed-decoder register substitutions examine which global image properties are accessible from registers and how their values affect reconstructed pixels. For generation, we train a multi-head attention pooling (MHAP) predictor to recover register features from sampled patches, keeping the pretrained diffusion model and its samples unchanged. On ImageNet-1K, our ViT-B decoders achieve mean reconstruction FIDs of 0.472 with DINOv2 and 0.442 with DINOv3, reductions of 9.5% and 5.2% relative to their respective patch-only baselines. The corresponding PSNRs are 19.73 and 20.54 dB. For generation, our DINOv2 models achieve a mean FID of 1.641, 9.8% below the patch-only baseline on identical sampled patches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.