acceptodds
Under review as a conference paper at ICLR 2027

Selective Attention for Training-Free Consistent Text-to-Image Generation

Abstract

Generating images with consistent object appearances across different scenes remains a fundamental challenge in text-to-image diffusion models. Existing approaches either require expensive fine-tuning for each subject or suffer from feature leakage when unconditionally sharing attention across images, where irrelevant backgrounds contaminate the target object's appearance. To address this, we first analyze the root cause of background leakage in naive cross-image self-attention fusion, identifying softmax normalization as the key factor that amplifies irrelevant token contributions. Grounded in this analysis, we propose SelectDiff, a training-free framework based on similarity-guided selective attention. Instead of fusing all cross-image tokens, we retain only the top-K most relevant connections per query token, naturally filtering background contamination while preserving subject identity. We further propose Next-Step Variance Ascent (NSVA), which dynamically adjusts K across denoising steps by maximizing between-class attention variance, and foreground-background separation to restrict cross-image attention exclusively to subject-relevant tokens. Experiments on the ConsiStory+ benchmark demonstrate that SelectDiff outperforms both fine-tuning-based and training-free methods in identity preservation and prompt alignment, while achieving superior inference efficiency. As a natural byproduct, our approach handles multi-subject scenarios without any specialized processing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.