SignRet: Sign Language Retrieval via Fine-Grained Gesture Descriptions and Ranking Refinement
Abstract
Sign language retrieval aims to establish bidirectional correspondence between sign language videos and natural language sentences. Existing methods typically match global video–text representations or aggregate local clip–word similarities. Global matching may overlook subtle gesture differences, while local matching relies on how well these details are encoded in the visual features. Moreover, both paradigms typically determine the final ranking from independently computed query–candidate similarities, without explicitly modeling the score and rank structure of the retrieved candidate list. Consequently, they may struggle to identify the correct match when multiple candidates exhibit similar visual patterns. To address these limitations, we propose SignRet, a sign language retrieval framework consisting of Sign Language Description Enhancement and Candidate Ranking Refinement. The former incorporates frame-level gesture descriptions into temporally corresponding visual features, enriching video representations with explicit fine-grained cues for more discriminative cross-modal matching. The latter revisits the top candidates obtained from the initial matching and models their contextual relationships to recalibrate candidate relevance, thereby improving the reliability of the final retrieval results. Experiments on PHOENIX-2014T, How2Sign, and CSL-Daily demonstrate that SignRet consistently improves both retrieval directions and achieves strong performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.