acceptodds
Under review as a conference paper at ICLR 2027

SignDiT: Hand-Faithful Pose-Guided Video Generation for Sign Language

Abstract

Pose-guided video generation can produce plausible human motion while corrupting the small hand articulations that determine meaning in sign language. This is not merely a resolution problem: conditioning must allocate geometric information across sources, adapt their reliability over space and time, and preserve discrete linguistic contrasts. We introduce SignDiT, a two-stage framework for these decisions. Stage 1 assigns image-plane structure to a 2D keypoint path and projection-ambiguous surface orientation, relative depth, and occlusion order to a second path, then modulates both with a part-conditioned continuous reliability field. Stage 2, Sign-DPO, changes one phonological factor at a time to construct counterfactual preference pairs and localizes credit to the affected region and frames. In self-reenactment and cross-identity evaluation on How2Sign, CSL-Daily, and Phoenix, SignDiT is best or tied best on 32 of 36 fidelity measures, reduces cross-identity FID by 25.1–41.2%, comes within 1.83 BLEU-4 points of real-video back-translation, and exceeds the strongest generated baseline by 3.22–3.79 points. The resulting design treats pose control as structured inference rather than feature concatenation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.