Beyond Scalar Assignments: Geometric Binary Quantization for Visual Tokenization
Abstract
Lookup-free quantization (LFQ) constructs large visual vocabularies compositionally from binary variables. Compared with conventional vector quantization, which relies on an explicit learned vector codebook and vector-level nearest-neighbor search, LFQ replaces this machinery with independent sign assignments, offering a more scalable discrete representation. However, this simplicity fixes the quantizer-side binary support to an axis-aligned hypercube. Although the encoder can adapt its representation toward this support, the quantizer geometry itself remains prescribed. To address this limitation, we introduce Geometric Binary Quantization (GBQ), which learns linear transformations to adapt the binary code geometry. However, directly applying a single transformation to the full binary latent would make exact nearest-code search prohibitively expensive. GBQ therefore factorizes the latent into small blocks and learns a local transformation within each block, enabling vector-level assignment locally while preserving exact and tractable assignment globally. Using the same tokenizer architecture and reconstruction objective on ImageNet-1K, GBQ achieves better reconstruction than WeTok with substantially fewer training-image exposures. Under matched-budget controlled training, the reconstruction advantage emerges early and persists throughout optimization, while GBQ maintains high code utilization without explicit entropy regularization. GBQ also retains competitive zero-shot reconstruction on out-of-domain data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.