AudioMouth: Articulation-Grounded Structured Generation for Speech-Driven Mouth Animation
Abstract
Speech-driven mouth animation aims to produce facial motion synchronized with speech, and existing audio-driven methods generally learn this mapping from acoustic features while leaving the articulatory structure of speech implicit in the waveform. Such implicit learning forces the generator to infer which mouth shape a sound requires from acoustics alone, and token-level objectives do not explicitly supervise the dynamics or geometry of complete motion sequences. In this paper, we aim to advance speech-driven mouth animation from implicit acoustic regression to explicit articulation-grounded generation. To this end, we propose AudioMouth, which formulates mouth animation as schema-constrained sequence generation with a multimodal language model. AudioMouth first segments long recordings into semantically coherent units sharing temporal boundaries across audio, transcript, phonemes, and target motion, so each phoneme sequence is paired with motion from the same segment. Conditioned on these units, AudioMouth generates fixed-schema ARKit coefficient sequences under explicit phoneme-to-articulation grounding, and reweights coefficient-value tokens to reduce the relative contribution of frequent near-zero targets to supervision. To supervise the trajectory as a whole, AudioMouth further refines the generator with Group Relative Policy Optimization, whose reward measures coefficient accuracy, temporal dynamics, and a simplified lip-geometry discrepancy that reads 3D mouth deformation out of a fixed blendshape basis without rendering or landmark detection. On our selected subject-disjoint BEAT split, using the provided linguistic annotations, AudioMouth improves every reported metric over the four evaluated baseline configurations, reducing Fr\'echet distance by 57% and coefficient MSE by 70% relative to the best baseline in each metric. Our ablations support the benefits of articulation guidance for distributional fidelity, balanced supervision for coefficient reconstruction, and geometric feedback for reduced across-script FD dispersion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.