acceptodds
Under review as a conference paper at ICLR 2027

Perturbing Text Attention for Guiding Diffusion Models

Abstract

Training-free guidance methods steer the sampling process of a diffusion model by extrapolating the prediction of the model away from a weak prediction. The weak prediction is obtained by perturbing intermediate computations of the model at inference time, such as self-attention weights or the arrangement of visual tokens. However, their performance gains over unguided sampling remain marginal for text-conditional generation. To understand why, we present an in-depth analysis of how diffusion models behave under conditions of varying complexity. Our analysis reveals that the primary bottleneck for text-conditional generation lies in aligning generated samples with the given condition, rather than in refining spatial details among visual tokens, which existing methods mainly focus on. Based on this, we introduce text attention guidance (TAG), a simple yet effective training-free guidance method that perturbs the interaction between visual and textual tokens. To this end, we propose to degrade the attention weights between the two modalities, yielding a weak prediction that particularly struggles to align samples with the given condition. This provides a targeted negative signal that mitigates the core bottleneck of text-conditional diffusion models. Extensive experiments on both U-Net and transformer-based models, including SDXL, SD3.5, Flux.1-dev, and Qwen-Image, demonstrate the effectiveness of our approach, which outperforms existing training-free guidance methods significantly and achieves results comparable to or even better than those of classifier-free guidance without any training. We will release the code upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.