Compatible Concept Geometry for Cross-Modal Compositional Generalization
Abstract
Vision-language dual encoders must transfer familiar concepts to combinations that are absent from training. Existing work relates compositional generalization to additive and orthogonal representations within each modality; whether independently structured image and text factors can be correctly matched on unseen combinations remains unclear. We study this cross-modal gap. We show that even when the two encoders separately exhibit additive and cross-factor orthogonal representations, incorrect factor correspondence can still cause compositional matching to fail, and we characterize the compatibility and identifiability conditions of cross-modal factors. To test this property, we develop order-wise geometric diagnostics and use Relative Orthogonal Concept Alignment (ROCA) as a structure-preserving intervention. Across three compositional datasets and eight pretrained models, text representations are generally closer to first-order additive structure, while corresponding image and text factors retain transferable, structured misalignment; first-order orthogonal alignment consistently improves unseen-combination performance, and compatibility weakens at higher response orders. Controlled experiments further show that first-order factor correspondence alone does not produce attribute-object binding, which requires interaction structure capable of distinguishing factor swaps. These results show that cross-modal compositional generalization is jointly constrained by within-modality factor structure, cross-modal factor compatibility, and concept binding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.