acceptodds
Under review as a conference paper at ICLR 2027

HUG: Learning and Transferring Visual Tokens for Unified Image Understanding and Generation

Abstract

We present HUG, a framework for unified image understanding and autoregres- sive image generation that combines visual tokenizer learning with vocabulary adaptation to a pretrained LLM. A shared vision encoder provides continuous features for understanding and discrete tokens for generation. HUG supplements post-quantization caption supervision and discrete-to-continuous alignment with direct caption and masked VAE-latent prediction before quantization. Controlled experiments show that direct pre-quantization supervision improves reconstruction fidelity and code-use diversity. Increasing the post-caption loss weight does not re- produce the pre-caption gains, and pre-quantization masked prediction outperforms its post-quantization counterpart in our setting. For vocabulary transfer, HUG builds on text-scale-aware normalization and uses Text-Aligned Normalization (TAN) with caption-supervised adaptation to initialize the LLM’s visual input and output weights. With the tokenizer fixed, the full TAN procedure outperforms the tested fixed and learned initialization alternatives without adding a downstream module. With a 0.6B Qwen3 language backbone, HUG achieves 0.85 on GenEval and 86.05 on DPG-Bench. Scaling to 4B yields 0.90 and 87.00, respectively, the highest overall scores on both benchmarks among the unified models evaluated, alongside competitive multimodal understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.