Fre-Flow: Offline Multiple Appropriate Facial Reaction Generation with Frequency-Aware Flow Matching
Abstract
Facial Reaction Generation (FRG) aims to synthesize contextually appropriate listener facial reactions conditioned on speaker behaviors in dyadic interactions. Since the same speaker behavior may elicit multiple appropriate reactions, FRG is inherently a one-to-many generation problem, motivating the Multiple Appropriate Facial Reaction Generation (MAFRG) task. However, existing MAFRG approaches typically employ unified spatio-temporal modeling despite the inherently multi-scale temporal structure of facial reactions. Such frequency-agnostic learning can cause cross-scale entanglement between global reaction trajectories and local variations, leading to unstable global dynamics and over-smoothed local details, while diffusion-based MAFRG approaches additionally incur costly iterative denoising for long-sequence generation. To address these challenges, we propose Fre-Flow, a frequency-aware Conditional Flow Matching (CFM) training strategy for offline MAFRG. Fre-Flow introduces two key components: (i) Frequency Flow Matching (FFM), which constructs frequency-aware spectral states and target velocities with adaptive noise scales, and (ii) Frequency Loss Optimization (FLO), which supervises velocity learning through frequency-weighted spectral matching and high-frequency energy regularization. During inference, Fre-Flow generates diverse listener facial reactions through efficient few-step ODE integration. Experiments on the MARS dataset show strong generation quality and inference efficiency, with a favorable balance among appropriateness, diversity, and structural fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.