SignDiT: Hand-Faithful Pose-Guided Video Generation for Sign Language
Abstract
Pose-guided video generation can produce plausible human motion while corrupting the small hand articulations that determine meaning in sign language. This is not merely a resolution problem: conditioning must allocate geometric information across sources, adapt their reliability over space and time, and preserve discrete linguistic contrasts. We introduce SignDiT, a two-stage framework for these decisions. Stage 1 assigns image-plane structure to a 2D keypoint path and projection-ambiguous surface orientation, relative depth, and occlusion order to a second path, then modulates both with a part-conditioned continuous reliability field. Stage 2, Sign-DPO, changes one phonological factor at a time to construct counterfactual preference pairs and localizes credit to the affected region and frames. In self-reenactment and cross-identity evaluation on How2Sign, CSL-Daily, and Phoenix, SignDiT is best or tied best on 32 of 36 fidelity measures, reduces cross-identity FID by 25.1–41.2%, comes within 1.83 BLEU-4 points of real-video back-translation, and exceeds the strongest generated baseline by 3.22–3.79 points. The resulting design treats pose control as structured inference rather than feature concatenation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.