RefineSLT: Learning What to Recover and When to Refine in Sign Language Translation
Abstract
Existing sign language translation (SLT) models can produce fluent sentences while omitting or mistranslating content essential to the signed message, such as entities, actions, or temporal information. Addressing these errors requires identifying potentially missing or mistranslated content and directing refinement toward its visual evidence. We introduce RefineSLT, a framework that couples semantic cue prediction with learned refinement decisions for pretrained SLT models. CueMiner derives semantic cue supervision from training translations, and CueGrounder learns to predict these cues and associate them with temporal visual evidence. CueRefiner combines cue confidence with estimated coverage in the initial translation to control bounded, temporally weighted feature updates. The pretrained decoder then generates the final translation from the refined features. The backbone encoder and decoder remain frozen, and inference requires neither reference translations nor ground-truth glosses. With enhancement modules trained for each backbone, the framework is designed to accommodate different architectures, translation capabilities, and sign languages, offering a common basis for enhancing existing translators and adapting to future models. An analysis of baseline outputs on PHOENIX-2014T illustrates how errors in a few content-bearing words or phrases can alter the intended message, motivating semantic cue learning and selective refinement. Experiments on PHOENIX-2014T, How2Sign, and Auslan-Daily demonstrate consistent BLEU-4 and ROUGE-L improvements across all evaluated backbones, including relative BLEU-4 gains of 12.1% on PHOENIX-2014T and 54.2% on Auslan-Daily Communication with CV-SLT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.