acceptodds
Under review as a conference paper at ICLR 2027

XLingoTalker: Cross-Lingual Personalized Speech-Driven 3D Facial Animation

Abstract

Cross-lingual personalized 3D facial animation aims to generate speech-synchronized facial motion in a new language while preserving a speaker's characteristic articulation patterns. Existing methods typically rely on language-matched speech-motion pairs and therefore cannot address cases where target-language motion is unavailable. We formulate this problem as weakly supervised cross-lingual personalized speech-driven 3D facial animation, where paired speech-motion data are available only in a source language, together with unpaired target-language speech from the same speaker. The key challenge is to separate speaker-specific motion style from language-dependent acoustic variation. We observe that audio-derived style representations vary substantially across languages, whereas motion-derived speaker representations remain more consistent. Motivated by this finding, we propose XLingoTalker, which grounds audio-derived style in a shared audio-motion space. A style-guided multi-head mixture-of-experts decoder maps the grounded representation to personalized facial dynamics, while a style-replacement consistency objective leverages unpaired target-language speech to preserve speaker-specific motion across languages. We further introduce XLingo-12, the first multilingual 3D talking-face dataset featuring controlled bilingual recordings from 12 speakers for fixed-identity cross-lingual evaluation. Experiments show that XLingoTalker better preserves personalized facial dynamics while maintaining speech-motion synchronization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.