PlanTalk: MLLM-Based Motion Planning with Hierarchical Audio Guidance for Talking Head Generation
Abstract
Audio-driven talking head generation is challenging because facial articulation is identity-dependent and fine-grained audio-visual correspondence is difficult to preserve across multi-stage generation. In this paper, we propose PlanTalk to exploit joint speech-identity context through an MLLM-based motion planner, together with hierarchical audio guidance. The motion planner jointly interprets the driving speech and speaker-specific visual characteristics to organize an identity-aware facial motion sequence. Hierarchical audio guidance then progressively strengthens audio-visual correspondence at different representation levels: audio-landmark fusion refines the alignment between facial geometry and speech articulation, while denoising-stage audio attention preserves speech-consistent facial dynamics during video synthesis. The proposed cross-modal planning strategy effectively captures identity-dependent articulation, while hierarchical audio guidance maintains fine-grained speech correspondence throughout the generation pipeline. Extensive experiments demonstrate that PlanTalk achieves state-of-the-art performance over existing methods in terms of lip synchronization and identity consistency while maintaining high visual fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.