UF-VSR: A Phonetic–Lexical Decomposition for Mandarin Visual Speech Recognition
Abstract
Mandarin visual speech recognition (VSR) must infer characters from lip motion that primarily constrains syllable identity: visually similar articulations can denote different syllables, and one Pinyin syllable can denote many characters. We present UF-VSR, a two-stage system that makes this mismatch explicit. A shared visual encoder with character and toneless-Pinyin decoders first produces complementary lexical and phonetic hypotheses; a 4B language model adapted to model-generated VSR errors then refines the character hypothesis conditioned on both. The probabilistic decomposition is a modeling interpretation rather than a claim of exact marginal inference: the implemented system uses one best hypothesis from each decoder and the language model does not observe video. On CMLR, refinement reduces character error rate (CER) from 11.781 to 10.166 (13.71% relative) for our primary upstream and from 8.021 to 7.603 (5.21% relative) for a stronger upstream. Fixed-track CNVSRC-Multi.Dev CER decreases from 30.87 for a character-only model to 28.39 for the complete pipeline. Mechanism-oriented analyses show where the system succeeds and fails: reliable Pinyin is associated with larger gains, identical-Pinyin errors require context, and refinement ceases to help when either phonetic or character evidence is severely corrupted. These results support a practical Mandarin-specific interface between visual and language models while delimiting, rather than overstating, the evidence for its individual components.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.