CIF: Commonality-Individuality Fusion for Creative Two-Object Composition
Abstract
Compositional object generation with text-to-image (T2I) models has attracted growing interest in creative design and content creation. However, existing object-fusion methods often rely on object-level representations or similarity measures, leading to superficial and structurally incoherent combinations. We propose Commonality–Individuality Fusion (CIF), a compositional generation framework that establishes shared structure before enhancing individual characteristics. First, Common–Individual Division (CID) decomposes two source objects into shared and individual semantic components. Then, Common Fusion (CF) integrates the shared components to construct a coherent structural scaffold and trains a fused diffusion model to generate the common foundation. Finally, Individual Enhancement (IE) incorporates the individual components and controlled perturbations to generate diverse fusion candidates, from which human preferences are used to optimize the fused diffusion model via direct preference optimization. Experiments show that consistently outperforms existing T2I and object-fusion methods, producing more coherent and distinctive compositional objects.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.