Taming Visual Tokenizer Friendly for Large Language Model
Abstract
Extending pretrained language models to visual understanding and generation can compromise their existing language and reasoning capabilities. We investigate which properties of visual tokenization affect this trade-off and find that the organization of visual latents matters beyond their standalone visual quality. We develop LMTok, a 1D visual tokenizer whose latent organization is shaped by the next-token prediction objective of a pretrained language model, and compare it with a spatial 2D tokenizer under matched token lengths, codebook sizes, and downstream training budgets. Despite comparable reconstruction, generation, and visual understanding performance, training with LMTok helps the LLM better retain its language capabilities than training with spatial 2D tokens. With Qwen3-8B as the backbone, LMTok improves over spatial 2D tokenization by 1.72 points on MMLU-Pro, 3.40 points on MATH-500, and 5.40 points on MBPP, while maintaining comparable visual performance under unified training. This trend remains consistent when scaling the language-model backbone from 4B to 8B. Comparisons with existing 1D tokenizers and controlled ablations identify the autoregressive shaping as a key contributor to this improved compatibility. Overall, our results show that visual tokenizers should be evaluated not only by standalone visual quality, but also by how their sequence organization interacts with the autoregressive prior. The code and model will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.