SignAlign: Structure-Aware Alignment for Cross-Modal Representation Learning in Sign Language Retrieval
Abstract
Sign language retrieval enables bidirectional search between natural-language text and signing videos based on their semantic correspondence. Existing methods compute video–text similarity by comparing global representations or aggregating local similarities between video tokens and text tokens. Despite their effectiveness, two limitations remain. First, contrastive training emphasizes annotated video–text associations without explicitly preserving semantic relations among texts. As text representations adapt to videos, their pretrained relative distances can change. If texts with different meanings become too close, a video associated with one may be incorrectly retrieved for the other. retrieved for the other. Second, local alignment computes weights independently for each text token. A video token with high similarity to several text tokens can increase the retrieval score through all of them, even when the video does not express the remaining content of the text. To address these limitations, we propose SignAlign, a structure-aware alignment framework for cross-modal representation learning. It preserves semantic relations among sentences while jointly aligning video tokens and text tokens. To preserve sentence relations during adaptation, Text Semantic Structure Alignment (TSSA) uses a frozen pretrained text encoder as a reference. It aligns normalized pairwise distances and triplet-wise angles among sentence embeddings with those derived from this reference. Complementing this sentence-level alignment, Cross-Modal Token Structure Alignment (CTSA) jointly determines token alignment weights through entropy-regularized optimal transport. Text tokens aligned with the same video token share its available weight, limiting repeated emphasis in retrieval scoring. The resulting joint alignment score is fused with independently weighted local similarity scores during training and inference. Experiments on three public sign language retrieval benchmarks demonstrate improved retrieval in both directions over strong baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.