Less Guessing! Hallucination-Aware Self-Critical Training for Sign Language Translation
Abstract
With the adoption of large-scale pretrained language models (LMs), Sign Language Translation (SLT) has achieved remarkable advances in generating fluent text from sign language videos. However, stronger language generation ability does not necessarily ensure faithfulness to the signed input, and contemporary SLT models may produce fluent translations containing visually unsupported actions or semantically distorted meanings. Despite the serious risks that hallucination poses to the reliability of SLT systems, this problem remains insufficiently studied. In this paper, we systematically investigate hallucination in SLT and characterize it into two forms: Action Hallucination at the visual-content level and Semantic Hallucination at the sentence-meaning level. Based on this characterization, we propose Hallucination-Aware Self-Critical Training (HalluSCT) for Sign Language Translation, a sequence-level optimization framework for hallucination mitigation. HalluSCT introduces a hallucination-aware reward consisting of two components: (i) a Visual Sensitivity Reward that favors translations exhibiting stronger dependence on salient visual evidence under counterfactual perturbation, (ii) a Semantic Consistency Reward that favors translations semantically consistent with the reference translation. The aggregated reward is optimized through self-critical sequence training, encouraging stronger reliance on visual evidence while maintaining semantic correctness. We further introduce SLT-oriented hallucination metrics for evaluating Action Hallucination and Semantic Hallucination. Extensive experiments on public SLT benchmarks, together with human and LLM-based evaluations, demonstrate that HalluSCT consistently reduces hallucination while improving translation quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.