acceptodds
Under review as a conference paper at ICLR 2027

Visual Tokenization with Latent Anchored Binary Quantization

Abstract

Discrete visual tokenizers are bottlenecked by both codebook scalability and quantization error. While lookup-free and finite scalar quantization improve scalability, they forgo the adaptive codebook updates used in vector quantization to reduce quantization error. This mismatch leads to degraded reconstruction and slower convergence in training. To address this, we revisit the two-latent variational formulation that interprets codebook learning as continuous-latent recovery within the ELBO. This interpretation motivates a learned recovery map for binary codes, however direct realization of this objective leads to collapsed representation when the map and its target latents are jointly optimized. We therefore introduce latent-anchored binary quantization (LABQ), a tokenizer that learns binary codes and a recovery map against the fixed representations of a pretrained VAE encoder while fine-tuning its decoder. LABQ trains in 30 epochs, using approximately 6.7x fewer tokenizer-training image exposures than BSQ based on the reported epoch budgets. Without adversarial supervision, some blurring and color-shift artifacts remain. We systematically compare PatchGAN, DINO, and pretrained VAE encoders as discriminators, finding that the SD3 VAE encoder, already attuned to the target latent geometry, yields the lowest rFID under the same ten-epoch adversarial fine-tuning schedule. With this design, LABQ achieves competitive results in discrete image reconstruction on ImageNet (rFID 0.29, LPIPS 0.038) and COCO (rFID 2.84, LPIPS 0.037), approaching its continuous VAE reference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.