acceptodds
Under review as a conference paper at ICLR 2027

StateWise: From Noise Awareness to State-Calibrated MLLM Conditioning for Unified Generation

Abstract

Unified multimodal models aim to improve image generation and editing through visual understanding. Connecting a pretrained multimodal large language model (MLLM) to a diffusion generator allows its understanding capabilities to guide image synthesis. However, reusing a fixed condition throughout denoising prevents the MLLM from adapting its guidance to the evolving latent. Adapting this condition requires identifying the current latent state, which may depart from the state expected at the input timestep during finite-step sampling.We introduce StateWise, a state-calibrated MLLM conditioning framework that reads noisy latents directly in the generator's native VAE space without an auxiliary visual encoder on this pathway. During training, we vary each latent's noise level while keeping its input timestep fixed, teaching the MLLM not only to perceive the latent's noise state but also to estimate how much noisier or cleaner it is than expected at that timestep. These estimates guide adjustments to the condition supplied to the diffusion generator at each denoising step, while an additional constraint discourages changes when the observed and expected noise levels agree. Using only publicly available training data, achieves 0.90 on GenEval and 77.4 on GenEval2, alongside competitive image editing. Code will be publicly released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.