A Language-Informed Loss for Visual Generation
Abstract
End-to-end visual generation with pixel-space loss has recently achieved competitive performance without the need for learned image tokenizers. In this work, we study whether language-informed supervision can provide a richer training signal than pixel loss for visual generation. We introduce **LILIE** (pronounced "lily"), a **L**anguage-**I**nformed and **L**ossless **I**mage **E**ncoder. LILIE is an invertible image transform trained such that a vision-language model can recover captions from noised features. The encoder uses a normalizing flow transformer trained jointly with a vision-language model via next token prediction. We then use the encoder to extract a simple language-informed loss for text-to-image generation. We train JiT models with LILIE loss, reducing gFD by 16 points (68.61 52.73) against a pixel-only baseline on the GPIC dataset. Our code, encoders, and generative models are included in the Supplementary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.