UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion
Abstract
We introduce UNITY, a Universal-to-Specialized adapter for efficient and scalable composite conditioning in diffusion-based image generation. Unlike prior approaches that train separate adapters for each condition, UNITY jointly learns shared semantics across multiple conditioning modalities and later specializes without modifying the architecture. The proposed two-stage training paradigm consists of a Universal Stage, which captures cross-modal joint representations across all conditioning types using half of the total training steps, followed by a Specialization Stage that refines modality-specific features with the remaining steps. At the core of UNITY are the Morphable Attention Flow (MAF) Network and Morph Wrapper modules, which enable channel-aware and spatially adaptive feature alignment through learnable flow fields and attention fusion. This constant-complexity formulation allows flexible operation under single or composite conditioning while significantly reducing inference latency and memory footprint. Extensive experiments across multiple datasets demonstrate that UNITY achieves state-of-the-art image fidelity and memory efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.