Unlocking Training-Free Continuous Attribute Control in Text-to-Image Diffusion Transformers via M-Space
Abstract
Text-to-image diffusion transformers (DiTs) have shown impressive capabilities in generating high-quality images, yet lack straightforward mechanisms to continuously adjust the strength of individual semantic attributes during image synthesis. To achieve such control, existing approaches often rely on training additional low-rank control adapters or optimizing semantic directions in textual embedding space, incurring extra computational costs. In this paper, we investigate whether the internal conditioning representations of pre-trained text-to-image DiTs can be directly used for continuous attribute control. Our analysis identifies a compact global modulation conditioning space, termed M-space, in which a shared conditioning representation is used to produce layer-wise modulation parameters in text-to-image DiTs. Attribute directions in M-space can be obtained by simply averaging representation differences across diverse pairs of text prompts that differ only in the target attribute. Linear traversal along these directions enables smooth and bidirectional attribute control, while their composition enables multiple attributes to be manipulated simultaneously. Furthermore, we introduce Self-Attention Map Blending (SAMB) and Attention-Guided Modulation Routing (AGMR) to retain attribute-irrelevant content during editing and enable instance- level attribute control, respectively. Extensive qualitative and quantitative experiments on text-to-image DiTs demonstrate that, without requiring any additional training, our method achieves performance comparable to existing state-of-the-art approaches in continuous attribute control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.