acceptodds
Under review as a conference paper at ICLR 2027

SignDict: Sign Language Production with a Compressed Pose Representation Dictionary and Adaptive Frame Weighting

Abstract

Sign Language Production (SLP) is a critical technology for breaking down communication barriers for people with hearing impairments and enabling interaction. Existing sign language production models are primarily divided into dictionary-based models and neural network-based models. Dictionary-based models offer high motion fidelity and good controllability, but they typically suffer from cumbersome generation processes, stiff gesture transitions, and high storage costs. Neural network-based models can generate more fluid sign language movements and possess strong generalization capabilities, but they are prone to pose distortion and semantic misalignment. This paper proposes SignDict, a sign language production model based on a feature dictionary. By combining the pose priors of traditional dictionaries with the continuous generation capabilities of neural networks, SignDict achieves high‑fidelity and fluent sign language pose generation. Unlike traditional dictionaries, which directly store complete discrete motion segments, SignDict constructs a pose feature dictionary that associates each gloss with a compact latent feature representation, representing sign language motions as latent feature vectors. During the generation process, the feature dictionary serves as a pose prior, guiding the model to generate pose sequences consistent with the semantic meaning. To achieve feature compression, this paper further designs an adaptive feature compression module that performs adaptive feature aggregation on pose segments of varying lengths corresponding to the same gloss, compressing them into a single latent feature vector. This significantly reduces the dictionary size while preserving motion information, achieving up to 1366‑fold compression in storage footprint compared with traditional motion dictionaries. Furthermore, to address the issue of varying importance of different pose frames in semantic representation, this paper proposes an adaptive frame weight loss function. While maintaining constraints on action transition frames, this function increases the weight of semantically critical frames, enabling the model to focus more on key actions and thereby further improving the accuracy and expressiveness of the generated poses. Experimental results on the two public datasets, PHOENIX14T and CSL-Daily, demonstrate that SignDict achieves competitive performance across multiple evaluation metrics and reaches state-of-the-art (SOTA) levels, validating the effectiveness of feature-dictionary-based sign language generation methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.