Hierarchical Cross-modal Representation Alignment for Text-to-image Diffusion Models
Abstract
Modern text-to-image diffusion models inherit rich compositional semantics from pretrained text encoders, yet their denoising objectives provide limited direct supervision for binding phrases to distinct image regions. Building on the success of representation alignment with external models, we harness the generator's own language representations as a source of hierarchical supervision. We introduce Hyperbolic Intermodal Projection (HIP), a lightweight method for cross-modal alignment in text-to-image diffusion models. HIP promotes fine-grained cross-modal alignment, which is stabilized by anchoring local representations to their global semantic context. To preserve the distinctiveness of concepts sharing the same anchor, we introduce a shared hyperbolic space whose geometry supports local separation within a coherent scene neighborhood. Experiments on SD3.5-M and FLUX.1-dev demonstrate improvements over baselines: HIP improves GenEval macro accuracy by and relative to the respective pretrained baselines, and achieves the highest scores across all six evaluated categories under the official T2I-CompBench evaluator on both backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.