acceptodds
Under review as a conference paper at ICLR 2027

Trust-to-Latent: Aligning Latent Generation with Multimodal Semantic Judgment

Abstract

The central challenge of long-prompt image generation is to realize entities, attributes, quantities, and relations jointly in a semantically coherent scene. Unified understanding–generation models offer a natural opportunity to address this challenge: their understanding capability can provide a semantic criterion for the content their generation capability seeks to realize. Yet judging completed images does not automatically enable understanding to assess the noisy representations encountered during generation. We introduce Trust-to-Latent, which connects these capabilities by extending reward-trained semantic judgment to intermediate generation states. A prompt-conditioned latent evidence interface makes multi-layer generation features accessible to a reader built on a frozen understanding backbone. Image-level preferences, transferred through shared-noise latent pairs, teach the reader to assess evolving visual evidence against the prompt. The complete evaluator then remains fixed while relative preferences among candidates from a common intermediate state guide generator learning. Understanding thus informs how the generator is trained to realize instructions, without reward evaluation or candidate selection at inference. Experiments across five public benchmarks show improved overall instruction following over the base model and reward-optimization baselines, with gains in attribute binding, spatial relations, and joint satisfaction of evaluated requirements. Interface ablations and paired error analyses further support this connection between semantic judgment and compositional generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.