SignXray: Plug-and-Play Sign Encoding in a Face-Hand-Body Multi-Prior Space
Abstract
Sign languages convey information through coordinated facial expressions, hand articulation, body motion, and their temporal interactions. Existing human-representation models expose complementary aspects of this signal through heterogeneous outputs that are difficult to reuse through a common video interface. We present **SignXray**, a plug-and-play sign video encoder learned in a unified Face-Hand-Body multi-prior space. During training, SignXray consolidates supervision from SMPLer-X, DWPose, MediaPipe, PrimeDepth, and Sapiens Segmentation into a compact, sign-aware representation and distills this heterogeneous prior knowledge into an RGB ViT. At inference time, the resulting encoder maps RGB video directly into the unified representation, giving downstream models access to the complementary information of all five prior formats through a single reusable feature interface. Experimental results demonstrate that SignXray achieves state-of-the-art performance on continuous SLR and translation tasks, with up to a 50× acceleration over the evaluated pixel-space baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.