acceptodds
Under review as a conference paper at ICLR 2027

Zero-Shot Multi-Concept Image Generation via Decoupled Cross and Self-Attention in MMDiT

Abstract

Generating high-fidelity images with multiple reference concepts is an important but challenging task. Existing methods either rely on additional training or suffer from identity blending, attribute leakage, and structural inconsistency in complex multi-reference scenarios. We propose Decoupled Cross- and Self-Attention MMDiT (Deco-DiT), a training-free framework for multi-concept image generation. Deco-DiT decouples multi-reference conditioning into Multi-Reference Cross-Attention (MR-CA) for localized identity routing and Spatial-Consistent Self-Attention (SC-SA) for reference-target structural modeling. MR-CA routes encoder-level visual features to LLM-parsed spatial regions, while SC-SA enables target tokens to access structural and appearance cues from reference latents. Together with a stage-wise denoising schedule, Deco-DiT improves identity preservation, text alignment, and compositional consistency across challenging multi-reference settings without additional training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.