StateWise: From Noise Awareness to State-Calibrated MLLM Conditioning for Unified Generation
Abstract
Unified multimodal models aim to improve image generation and editing through visual understanding. Connecting a pretrained multimodal large language model (MLLM) to a diffusion generator allows its understanding capabilities to guide image synthesis. However, reusing a fixed condition throughout denoising prevents the MLLM from adapting its guidance to the evolving latent. Adapting this condition requires identifying the current latent state, which may depart from the state expected at the input timestep during finite-step sampling.We introduce StateWise, a state-calibrated MLLM conditioning framework that reads noisy latents directly in the generator's native VAE space without an auxiliary visual encoder on this pathway. During training, we vary each latent's noise level while keeping its input timestep fixed, teaching the MLLM not only to perceive the latent's noise state but also to estimate how much noisier or cleaner it is than expected at that timestep. These estimates guide adjustments to the condition supplied to the diffusion generator at each denoising step, while an additional constraint discourages changes when the observed and expected noise levels agree. Using only publicly available training data, achieves 0.90 on GenEval and 77.4 on GenEval2, alongside competitive image editing. Code will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.