MoLaGen: Multimodal Conditional Layout Control for Image Generation
Abstract
In the layout-to-image generation task, instances are typically guided by local prompts and spatial coordinates. Since text-based conditions alone are often insufficient for precise instance-level specification, incorporating a reference image for each instance enables more accurate and fine-grained conditional description. To achieve such multimodal conditional layout control, it is essential to encode and fuse visual, textual and spatial modalities, thereby enabling effective instance-level conditioning. Most existing approaches rely on separate encoders and fusion modules for multimodal conditioning, which may lead to representational discrepancies with the base generative model. Instead, we leverage the MM-DiT model itself to encode local prompts and reference images, and concatenate the resulting tokens as unified generation conditions. For the spatial modality, we propose a spatially-aware CondLoRA fine-tuning network that introduces spatial-specific mapping weights and employs token-dependent routing to activate different CondLoRA branches for spatial control. Finally, we propose the **M**ultim**o**dal Conditional **La**yout Controllable **Gen**erative Network (MoLaGen). In addition, we introduce a new dataset construction pipeline for synthetic data generation, instance-level annotation, and data filtering. Qualitative and quantitative experiments demonstrate the effectiveness of our method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.