Bind-and-Route: Explicit Reference Binding and Attention Routing for Multi-Subject Personalized Generation
Abstract
Multi-subject personalized generation aims to synthesize multiple customized subjects in a single image while preserving both textual controllability and reference appearance fidelity. However, we observe that existing personalization methods degrade substantially as the number of reference subjects increases, leading to reduced text alignment and weakened appearance consistency. By analyzing multi-modal attention patterns, we find that such degradation is closely related to Text-Reference Misalignment and Subject Confusion: the text of a subject and its corresponding reference condition may attend to different generation regions, while the target generation regions of a subject may attend to mismatched reference conditions. To address these issues, we propose Bind-and-Route, a framework for robust multi-subject personalization. First, we establish explicit text-reference associations through placeholder-reference contrastive learning and bidirectional binding resamplers, which mutually inject the associated textual and reference information to facilitate binding in multi-modal attention. Second, we introduce a lightweight attention router, which uses image-text attention as a localization prior to route image tokens toward matched reference conditions. Experiments demonstrate that our method can achieve strong text alignment and reference appearance fidelity, offering an effective solution under unified multi-modal conditioning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.