LEMO: LLM-Guided Fine-Grained Emotion Editing for Speech-Driven Facial Motion Generation
Abstract
Although speech-driven facial motion generation has advanced, faithfully synthesizing expressions that match emotional prosody remains a challenging task. This is mainly due to the close coupling of emotional cues and phonetic content within audio signals, which results in learned representations where these aspects are entangled. From our analysis, we observe two prevalent failure modes in current models: Expression Attenuation, where facial animations lack the intended emotional intensity even when the audio input is expressive, and Emotion Mismatch, where the model reproduces emotions from phonetically similar training samples instead of the user's desired emotion. Rather than addressing these limitations solely during generation, we propose LEMO, a motion-level emotion editing framework. LEMO decomposes facial motion into upper-face movement, phonetically driven neutral lip motion, and expressive lip offsets, using a phone-driven decomposition approach. To support flexible, text-based editing, expressive offsets are discretized as structured region-state descriptors, achieved by spatial and temporal quantization of blendshape activation. This tokenized interface enables a large language model to interpret natural language instructions and deliver region- and time-specific expression changes. A diffusion-based temporal inpainter, together with a region-aware ControlNet, then applies these modifications additively, preserving lip synchronization. Quantitative experiments on the MEAD dataset show that LEMO sets a new benchmark for both expression fidelity and motion naturalness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.