Trust-to-Latent: Aligning Latent Generation with Multimodal Semantic Judgment
Abstract
The central challenge of long-prompt image generation is to realize entities, attributes, quantities, and relations jointly in a semantically coherent scene. Unified understanding–generation models offer a natural opportunity to address this challenge: their understanding capability can provide a semantic criterion for the content their generation capability seeks to realize. Yet judging completed images does not automatically enable understanding to assess the noisy representations encountered during generation. We introduce Trust-to-Latent, which connects these capabilities by extending reward-trained semantic judgment to intermediate generation states. A prompt-conditioned latent evidence interface makes multi-layer generation features accessible to a reader built on a frozen understanding backbone. Image-level preferences, transferred through shared-noise latent pairs, teach the reader to assess evolving visual evidence against the prompt. The complete evaluator then remains fixed while relative preferences among candidates from a common intermediate state guide generator learning. Understanding thus informs how the generator is trained to realize instructions, without reward evaluation or candidate selection at inference. Experiments across five public benchmarks show improved overall instruction following over the base model and reward-optimization baselines, with gains in attribute binding, spatial relations, and joint satisfaction of evaluated requirements. Interface ablations and paired error analyses further support this connection between semantic judgment and compositional generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.