Semantic-REACT: Temporally Grounded Linguistic Conditioning for Multiple Appropriate Facial Reaction Generation
Abstract
In dyadic human-human interactions, linguistic semantics help human listeners interpret the speaker's verbal behaviour and determine when to express appropriate facial reactions (AFRs). Existing automatic Multiple Appropriate Facial Reaction Generation (MAFRG) methods typically generate listener facial reactions conditioned on speaker audio-visual behaviour, where encoded speech representations primarily capture acoustic cues, without explicitly modelling word-level semantics or when those semantics are relevant to the listener’s response. Consequently, the generated AFRs are usually appear plausible in motion while remaining semantically mismatched or temporally misaligned with the interaction context. To address this limitation, we propose Semantic-REACT, the first MAFRG solution that specifically encodes speaker's verbal transcripts and aligns the resulting linguistic features with the generated listener facial motion timeline. A gated adapter then injects these aligned linguistic features into the diffusion denoiser. This linguistic conditioning provides temporally aligned guidance while preserving the original speaker audio-visual conditions and the diversity of the generated listener AFRs. Experiments on the official REACT 2025 MARS benchmark show that Semantic-REACT improved the appropriateness of the generated listener facial reaction compared to Trans-VAE, PerFRDiff, and REGNN while preserving decent diversity. Matched control experiments confirmed the importance of temporal alignment for linguistic conditioning. Qualitative analysis and human evaluations further indicate that the generated listener facial reactions exhibit temporal coherence, semantic appropriateness, and favourable perceptual quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.