Resolving Spatial and Semantic Conflicts in Multi-LoRA Composition via Softmax Routing and Cross-Adapter Attention Conditioning
Abstract
Multi-LoRA composition activates several Low-Rank Adaptation (LoRA) modules simultaneously so that a single diffusion forward pass generates an image that embodies multiple personalized concepts at once. Existing methods either merge adapter weight deltas uniformly corrupting concepts that share spatial territory—or route noise predictions via hard per-pixel argmax, which discards non-dominant LoRA contributions and collapses identity consistency on overlapping pairs. We propose two complementary, training-free mechanisms. SR-LoRA (Spatially-Routed LoRA) replaces hard argmax routing with temperature-controlled softmax masks over per-LoRA cross-attention maps, smoothly blending noise predictions at each spatial position. CLAC (Cross-LoRA Attention Conditioning) runs a two-pass UNet forward where the anchor LoRA's cross-attention keys and values are blended into the secondary LoRA's attention, suppressing appearance leakage; an IoU-based overlap detector activates this conditioning adaptively across all pair types. On the ComposLoRA benchmark (240 compositions per domain, 10 seeds), SR-LoRA and FreeFuse reach equivalent CLIP-T (+4.3% over LoRA-C), while CLAC improves CLIP-I by +7.6pp over FreeFuse on reality (0.915 vs. 0.839) and +0.9pp on anime. The combination (Adaptive Fusion) achieves the highest CLIP-I on anime (0.875), validated with per-pair-type breakdowns, ablations, and N=3 triplet experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.