Rethinking Domain Generalization in Lip Reading: A Dependency Reliability Perspective
Abstract
Lip reading, which recovers spoken content from lip dynamics, has made substantial progress on standard benchmarks. However, these gains do not translate to real-world settings, where shifts in speakers, recording conditions, and linguistic distributions can sharply degrade performance. Existing domain generalization methods tackle these shifts primarily by learning domain-invariant visual representations, yet they overlook that lip-reading models inherently rely on linguistic context to compensate for visual ambiguity, thereby internalizing strong context-word dependencies. When these dependencies become unreliable across domains, such source-specific cues turn into misleading priors, causing severe generalization failures. To address this challenge, we propose Selective Dependency Gradient Projection (SDGP), a model-agnostic strategy that explicitly regulates non-transferable linguistic dependencies during optimization. Specifically, SDGP identifies dependencies exhibiting reliability mismatch across episodic domain splits, and quantifies model reliance on them via source-favored incorrect candidates. Differentiating this score yields a reliance gradient that captures the optimization direction amplifying these shortcuts. To suppress such reinforcement, SDGP projects the recognition gradient onto an adaptive half-space that imposes a first-order suppression constraint. The resulting convex subproblem admits a closed-form solution. To systematically study this mechanism, we further introduce LipWild-G7, a seven-genre Chinese benchmark designed for controlled cross-domain evaluation under joint visual and linguistic shifts. Extensive experiments on Chinese and English benchmarks show that SDGP consistently improves unseen domain recognition, highlighting dependency reliability in cross-domain lip reading.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.