When Frequency Meets Sign: Exploiting Complementary Signals for Vision-Only Sign Language Sentence Segmentation
Abstract
Most existing sign language translation (SLT) systems are developed for sentence-level video clips, limiting their ability to handle long videos containing multiple sentences in real-world scenarios. To bridge this gap, we introduce the Vision-only Sign Language Sentence Segmentation (VSLSS) setting, which partitions long multi-sentence videos into sentence-level clips solely from visual signals. As a preprocessing step, VSLSS allows existing sentence-level SLT models to process multi-sentence videos without retraining on multi-sentence data. However, sign language often lacks explicit motion pauses between sentences, making accurate sentence boundary localization from visual information alone challenging. To address this challenge, we propose FVS, a frequency-aware vision-only sign language sentence segmentation framework. To better exploit visual cues for sentence boundary localization, FVS introduces a Skeletal Motion Cepstral Coefficient (SMCC) extractor to capture additional high-frequency motion information beyond pose features. The pose and motion-frequency features are processed by two parallel branches and adaptively fused through a learnable Gated Fusion module. To alleviate under-segmentation, over-segmentation, and fragmented predictions, we further design Prior-Guided Losses (PGL). Extensive experiments demonstrate the effectiveness of FVS, and downstream experiments further validate its utility for sentence-level SLT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.