acceptodds
Under review as a conference paper at ICLR 2027

Real-Time Sign Language Translation by Prototype-Assisted Causal Streaming

Abstract

Sign language translation (SLT) for live captioning must produce text while the signer is still signing: each decision may use only frames already observed, and text shown to the user cannot be revised. Decisions are therefore made from short windows in which a sign may be only partly visible and pose detectors lose joints, so window-level recognition must be stabilised before it is committed. We propose a real-time, strictly causal, pose-based SLT pipeline with three components. A confidence-aware masked pose autoencoder uses detector confidence to gate unreliable joints and unreliable reconstruction targets. A part-specific vector-quantised tokenizer, with separate body, hand, and face codebooks trained by video-level contrast, block-to-video alignment, and angular-margin and centre objectives, maps short windows to discriminative sign units. A causal streaming decoder retrieves the nearest sign prototype for each window, aggregates retrievals by past-only voting, and passes the gloss stream through a causal boundary head to a Transformer translator; it reads no future frames and commits each output once. We further formalise streaming SLT through window end-times, zero lookahead, and commit-on-emit, and report time-to-first-token, token and end-to-end latency, and the quality and coverage of committed text. On Phoenix-2014, Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL, the system improves BLEU-4 over the strongest online baseline by 5.55 on Phoenix-2014T and 6.29 on CSL-Daily, keeps a 2.59 BLEU-4 lead at matched time-to-first-token, and has committed 84% of its final output after 80% of the input. Controlled ablations isolate each component, the train-side word segmentation used to build the prototype bank, and the pose detector.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.