acceptodds
Under review as a conference paper at ICLR 2027

Spatial-Query Multimodal Routing for Audio-Driven 3D Gaussian Talking Heads

Abstract

Audio-driven 3D talking heads require a representation that maps global speech signals to localized facial motion while preserving identity and rendering fidelity. Although spatial-query conditioning has been adopted in dynamic 3D Gaussian talking heads, its interaction with local reconstruction objectives and its consistency across identities remain underexplored. Building on GaussianTalker, we study a dynamic 3D Gaussian talking-head framework based on Gaussian-wise spatial-query conditioning. Multi-scale planar features encode the spatial context of each primitive, while a Spatial-Audio Cross-Attention (SACA) module uses these features as queries to route information from audio, eye state, camera pose, and a learnable null condition. The resulting Gaussian-specific conditioning predicts residual updates to position, scale, rotation, and appearance, enabling speech-driven deformation to vary across facial regions. A progressive training schedule first optimizes global reconstruction and Gaussian structure and then refines local non-rigid deformation with lip-region reconstruction and foreground-mask supervision. Under an aligned-frame evaluation protocol across three identities, the full configuration achieves the highest macro-averaged PSNR and SSIM among the compared configurations, reaching 33.52 dB and 0.936, respectively. In a 36-run 2-by-2 factorial study of SACA and lip supervision, nine matched Full/No-SACA comparisons show that the full model reduces mean LSE-D by 0.411 and increases mean LSE-C by 0.744, with a mean PSNR decrease of 0.181 dB. Across six directed source-to-target identity pairs, ArcFace cosine similarity ranges from 0.910 to 0.947. These results show that the empirical benefit of the complete spatial-query fusion configuration depends on local supervision and identity, improving average audiovisual synchronization under lip supervision while incurring a measurable reconstruction trade-off.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.