acceptodds
Under review as a conference paper at ICLR 2027

Text-to-Image Generative Modeling Without Language Pretraining

Abstract

Modern text-to-image (T2I) systems use internet-scale language pretraining for semantic conditioning. We ask whether language pretraining is necessary for strong T2I generation. We introduce **Xanthus**, a family of early-fusion multimodal autoregressive models pretrained entirely from scratch on GPIC, a dataset of 100M image-text pairs. **Xanthus-0.52B** outperforms LlamaGen-XL/8 (**48.0** vs. 63.5 gFD) with **18.2×** less pretraining compute. Further, **Xanthus-1.6B** outperforms JiT-T2I-XXL/8++ (**35.3** vs. 43.2 gFD) with **275×** less pretraining compute. Controlled data composition and scheduling experiments show that external text-only pretraining does not improve T2I generation. Under the same pretraining budget, **Xanthus** also learns image captioning from the paired image-text data while maintaining competitive T2I generation quality. We show that strong T2I systems can be trained from scratch without internet-scale language pretraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.