FLAT-SID: Resampling Image and Text into Flexible-Length Aligned Semantic IDs via Learned Step Quantization
Abstract
Representing visual content as discrete 1D tokens enables efficient storage, indexing, and seamless integration with generative models. However, existing 1D visual tokenizers such as TiTok and FlexTok prioritize pixel reconstruction over language alignment and rely on training a downstream model for cross-modal generation tasks. In this work, we introduce FLAT-SID, a discrete representation pre-training framework that jointly learns Flexible-Length, Aligned, Transmodal semantic ID tokens and their decoders, unifying cross-modal retrieval and generation (T2I and I2T) within a single architecture. A core challenge in end-to-end discrete training is to disentangle representation learning and quantization. We address this by applying quantization-aware training (QAT) to continuous latents. Specifically, we use Learned Step Quantization (LSQ) from the QAT literature, learning adaptive dimension-wise step sizes without the rigid grid constraints of lookup-free vector quantizers like FSQ or LFQ. Building on LSQ, FLAT-SID supports flexible-length semantic ID outputs for sparse retrieval and cross-modal generation, allowing for a tunable trade-off between performance and storage efficiency. The pre-trained FLAT-SID achieves 75.4 GenEval scores and, with lightweight fine-tuning, delivers competitive performance among state-of-the-art sparse methods: T2I generation with 81.2 GenEval score, I2T generation on MSCOCO with 139.4 CIDEr score, and T2I retrieval with Recall@5 70.9 and I2T 83.9 both on MSCOCO. FLAT-SID demonstrates the effectiveness of QAT for discrete semantic representations with a unified contrastive-generative framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.