TIDE: TEXT INTERVENTION FOR DIVERSIFIED EX- PLORATION IN DIFFUSION MODELS
Abstract
Accelerated text-to-image models, including distilled generators such as Z-Image-Turbo and FLUX.2-Klein, can produce high-quality and semantically aligned images in only a few denoising steps. However, repeated sampling of the same prompt often yields highly similar compositions and styles despite different initial noise seeds. We refer to this reduced sensitivity to seed variation as noise neglect, which manifests as a near one-to-one prompt-to-image mapping and severe within-prompt diversity collapse. Existing training-free approaches improve diversity through various inference-time interventions, but often trade off generation quality, efficiency, or architectural generality. In this work, we propose to apply intervention in the text embedding space as a novel framework for achieving rich diversity in diffusion models. By explicitly decomposing the text representation with Singular Value Decomposition (SVD), we propose a new spherical linear interpolation method for the decomposed right-singular vectors of the text conditioning, which is enriched with emergent image structure. However, such manipulation of the prompt embedding alone does not explicitly enforce separation among different seed trajectories. We also propose a latent-space repulsion method to coordinate trajectories sharing a prompt. By measuring similarities across multiple aspects—such as spatial and channel statistics—we derive a repulsive loss to explicitly push the latents apart and enhance diversity. Extensive experiments show that our method effectively recovers the “one-to-many” generative capability, significantly enhancing visual diversity while maintaining fidelity. Our code will be open-sourced.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.