GAMUT: Unified Visual Representation for Generation and Multimodal Alignment
Abstract
Visual representation learning has evolved through largely separate paradigms, including self-supervised learning, multimodal alignment, and visual tokenization. Despite substantial progress within each paradigm, this separation is increasingly at odds with the demands of modern vision and multimodal systems, which require diverse information to support open-ended tasks rather than representations specialized for a single capability. To address this issue, we introduce GAMUT, which learns a unified visual representation and brings these paradigms into a single joint pretraining framework. We develop three key techniques to ensure stable and efficient joint pretraining. First, we construct a three-stage training pipeline that progressively shapes the representation and avoid the prohibitive computational cost of prolonged multi-objective training. Second, we propose the latent channel regularization, which effectively mitigates the channel-wise energy distribution imbalance and prevents representation collapse Third, we introduce latent interpolation decoding to further improve the suitability of such representation for generative tasks. Together, these designs enable GAMUT to capture high-level semantics while preserving fine-grained visual details. As a result, GAMUT performs strongly across a diverse range of visual tasks, outperforming existing unified tokenizers and task-specific baselines on multiple benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.