acceptodds
Under review as a conference paper at ICLR 2027

Relaying Understanding to Generation in Unified Multimodal Models

Abstract

Unified multimodal models bring visual understanding and generation into a single architecture, yet sharing a backbone alone does not guarantee that knowledge acquired for understanding can be effectively utilized by the generator. A central challenge is therefore not only what the model understands, but also how understanding-side representations are exposed in a form that the generation pathway can readily consume. We address this problem by introducing an explicit understanding-to-generation interface into the topology of unified multimodal models. Specifically, understanding- and generation-oriented computation are separated, while a compact set of learnable relay tokens is inserted between them. The understanding pathway transforms input semantics into these intermediate states, and the generation pathway is constrained to condition on them rather than directly accessing the original input representation. This topology encourages the understanding pathway to organize semantic information into a generation-compatible representation, while providing the generator with a compact and structured conditioning interface. Despite its simplicity, the proposed design consistently improves image generation, with particularly strong gains on compositional semantics such as spatial relations, attribute binding, and actions. Our results suggest that explicitly structuring the topology of information flow between understanding and generation is an effective way to improve how unified multimodal models translate semantic understanding into visual generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.