Zero-Shot Multi-Concept Image Generation via Decoupled Cross and Self-Attention in MMDiT
Abstract
Generating high-fidelity images with multiple reference concepts is an important but challenging task. Existing methods either rely on additional training or suffer from identity blending, attribute leakage, and structural inconsistency in complex multi-reference scenarios. We propose Decoupled Cross- and Self-Attention MMDiT (Deco-DiT), a training-free framework for multi-concept image generation. Deco-DiT decouples multi-reference conditioning into Multi-Reference Cross-Attention (MR-CA) for localized identity routing and Spatial-Consistent Self-Attention (SC-SA) for reference-target structural modeling. MR-CA routes encoder-level visual features to LLM-parsed spatial regions, while SC-SA enables target tokens to access structural and appearance cues from reference latents. Together with a stage-wise denoising schedule, Deco-DiT improves identity preservation, text alignment, and compositional consistency across challenging multi-reference settings without additional training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.