TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading
Abstract
Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and motion sensitivity of event cameras. However, existing methods typically perform spatial compression before sufficient temporal modeling, which may suppress sparse and localized motion trajectories that are crucial for distinguishing similar lip movements. Moreover, most current approaches optimize temporal representations mainly at the word-classification level, leaving the underlying articulatory structure weakly constrained. To address these limitations, we propose a temporal aggregation framework that preserves localized motion trajectories and introduces viseme-level sequence guidance. First, we introduce Trajectory-Aware Differential Aggregation (TDA), which performs local temporal modeling at each spatial location before adaptive spatial aggregation. Second, we propose Viseme-Guided Aggregation (VGA), which introduces viseme-level sequence supervision to guide temporal aggregation for word recognition. Third, we employ EMA teacher-student consistency to regularize training and improve generalization. Extensive experiments on DVS-Lip and DVS-LRW100 show that the proposed method establishes new best results under the corresponding evaluation settings, while qualitative analyses further reveal that the learned temporal representations capture meaningful viseme-aware structure from event streams.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.