acceptodds
Under review as a conference paper at ICLR 2027

CoRe: Jointly Learning Contextual Representations for Understanding and Generation

Abstract

Understanding and generation place distinct demands on learned representations, yet these demands need not conflict. We introduce CoRe which jointly learns contextual representations that support understanding, provide conditioning, and serve as the space for generation. Both images and text are generated through continuous flows in the representations used for understanding. Trained from scratch on CC12M, the 383M-parameter model matches same-data contrastive controls on global understanding and reaches COCO-30K FID 10.10, comparable to much larger generators with pretrained components. Stronger language-pretrained systems retain an advantage in prompt following and caption fidelity. Matched interventions show that generating directly in the representation used for understanding improves compositional generation over a separate projected space, even as the latter reconstructs more accurately. A more flexible readout also recovers most of the dense-prediction gap. Together, these results show that understanding and generation can coexist in one contextual representation, with their interaction governed largely by how information is retained, organized, and made accessible.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.