acceptodds
Under review as a conference paper at ICLR 2027

A Unified View of Jointly Evolving Representations and Images in Diffusion Models

Abstract

Diffusion models increasingly use pre-trained visual representations to guide image generation, enhancing convergence and generation quality. Existing methods, however, mostly couple representations and images at two endpoint timesteps: RCG first synthesizes a clean representation at and then uses it as a static condition, while ReDi and REG jointly denoise the representation and image from pure noise at . In this paper, we provide a unified view of understanding such contradictions and show that both endpoints are suboptimal. Intuitively, the former prevents representations from receiving image context, while the latter forces image denoising to rely on highly noisy representations. We formalize this trade-off as a minimization problem, where these two setbacks combined, coincide with the divergence between actual diffusion denoising and a hypothetically ideal one, with clean representations or images as conditions. Empirical probing and theoretical analysis show that the optimum lies at neither extreme endpoint, but at an intermediate coupling timestep . Motivated by this insight, we introduce **JERI** (Jointly Evolving Representations and Images), which first denoises representations to this intermediate state and then jointly evolves representations and images. Experiments show that JERI outperforms the extreme designs in prior works across conditional and unconditional generation, model scales, and representation spaces.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.