FreeID: Calibrating Attention for Identity Preservation in Image-conditioned T2I Models
Abstract
Text-to-Image (T2I) diffusion models have driven rapid progress in personalized image generation, yet preserving subject identity from a reference image remains a central challenge. Conventional pipelines add identity through per-subject optimization or learned identity-conditioning modules trained on large-scale data. More recently, image-conditioned T2I (IT2I) models condition on reference images natively, processing text, reference images, and generation tokens within a unified attention framework. Yet this native conditioning still fails in demanding cases such as cross-view generation or small-scale subjects. Our analysis traces this failure to a specific mechanistic gap: IT2I models identify which reference tokens to attend to, but lack a mechanism to calibrate how much attention identity-critical features should receive, under-weighting them relative to peripheral ones; boosting facial attention alone recovers identity. Motivated by findings from human face recognition that internal facial regions carry disproportionate identity weight, we propose FreeID, a training-free framework that closes this gap at inference time. FreeID exploits the model's own attention responses to match reference facial keypoint tokens to their corresponding generation tokens, then injects sparse attention biases that amplify these identity-critical links, with magnitudes set by a closed-form bound on attention-output drift that adapts per token and step. Applied to state-of-the-art IT2I models, FreeID improves face similarity by up to 17.2% with negligible quality degradation, and further improves learned identity modules and image-to-video (I2V) models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.