Sparse-Landmark Non-Rigid Alignment of Spectral Point Sets for Cross-Modal Retrieval
Abstract
Cross-modal retrieval typically learns shared representations from abundant paired data. We study a complementary regime in which unimodal representations are independently constructed and kept fixed, while reliable cross-modal correspondences are scarce. The resulting challenge is therefore not representation learning, but sparse registration between two latent spaces. Our key observation is that independently learned spectral spaces can retain non-random semantic neighborhood agreement despite occupying different coordinate frames. This motivates modeling their mismatch as a dominant global coordinate discrepancy followed by residual, position-dependent non-rigid deformation. We accordingly formulate sparse cross-modal alignment as a global-to-local geometric registration problem and propose Sparse-Landmark Non-Rigid Alignment (SLNA). Using verified pairs as sparse landmarks, SLNA first recovers a global distance-preserving coordinate relation and then allocates nonlinear capacity only to the remaining residuals through regularized corrections in complementary coordinate and angular geometries. Experiments on Flickr30k and MSCOCO show that SLNA consistently improves over rigid Procrustes alignment and also outperforms a direct RBF mapping under matched model-selection settings. Structural diagnostics reveal statistically significant cross-modal neighborhood agreement, while controlled-deformation experiments show that residual modeling recovers local distortions beyond a single global transformation. These results suggest that sparse cross-modal alignment benefits from treating global coordinate recovery and local deformation modeling as distinct estimation problems, rather than learning a single flexible mapping from sparse pairs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.