ATLAS: Timed Lexical Anchors for Sign Language Production at Scale
Abstract
Glosses, ordered transcriptions of signs as written words, considerably improve sign language production (SLP) when used as a condition, but require expert annotation that large signing corpora lack. Glosses derived from the text alone do not reproduce this gain, as they carry no information about when each sign occurs. We instead extract timed lexical anchors from the signing itself with a sign-identity model trained on isolated sign dictionaries, and combine them with the sentence as a hybrid condition: the anchors supply the timing of the signs they cover and the sentence supplies the remaining content. At inference, the anchors are predicted from the sentence alone. The scheme requires no gloss annotation of the training corpus and therefore extends to large unannotated datasets. We extract and release SMPL-X parameters for more than 1,700 hours of video from YouTube-SL-25, and train ATLAS, a continuous-latent flow-matching generator, on more than 600 hours of ASL signing, substantially more than used in prior work. The anchors improve back-translation at every data scale we test: an anchored model trained on a fraction of the data surpasses a sentence-only model trained on all of it, and relative to each system’s ceiling, ATLAS reaches more than twice the back-translation score of the strongest baseline. Finally, we audit the SLP evaluation protocol used in the literature with content-free baselines, outputs that carry no signing by construction. Reference-based distance metrics favour low-amplitude motion, to the extent that such outputs can score higher than state-of-the-art models. We propose a protocol that reports the model’s performance grounded in these baselines and measures content with a public back-translation model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.