MNCFR: Embodied Simulation-Inspired Reaction Formation and Memory-Augmented Regulation for Multiple Appropriate Facial Reaction Generation
Abstract
In dyadic human–human interactions, observing a speaker's emotional expressions may elicit emotional contagion and facial mimicry in the listener, with embodied simulation considered a potential underlying mechanism. Existing Multiple Appropriate Facial Reaction Generation (MAFRG) methods typically learn direct mappings from the given speaker behavior to multiple appropriate facial reactions (AFRs), without explicitly modeling the perception-to-response transformation. Inspired by embodied simulation, we propose Multimodal Neural Context Formation and Regulation (MNCFR), a novel diffusion framework comprising two complementary components: the Reaction Formation Module (RFM) and the Memory-Augmented Regulation Module (MAR). Specifically, RFM uses listener-specific information derived from listener reaction history to transform the speaker's currently perceived multimodal behavior into an intermediate listener-side reaction representation and then into a generation condition for the current window. This explicit speaker-to-listener transformation models the perception-to-response pathway proposed in embodied simulation. The same intermediate representation is also fed to an auxiliary EEG prediction head, which predicts segment-aligned listener EEG and provides an additional listener-side training signal. Meanwhile, MAR applies diffusion-step-dependent gating to a bounded-memory readout of recent speaker representations and injects the resulting context into the RFM condition, thereby promoting smoother transitions between the generated listener facial reactions across consecutive time windows. By separating current reaction formation from recent-context regulation, MNCFR organizes complementary information across temporal scales to support coherent reaction generation. Experiments under the official offline and online protocols on MARS show that MNCFR achieves better overall performance than reported learning-based methods. Ablation and temporal analyses demonstrate the effectiveness of RFM and MAR, while EEG prediction analysis further shows that higher prediction quality based on RFM's intermediate listener-side reaction representation is associated with larger generation gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.