acceptodds
Under review as a conference paper at ICLR 2027

How do diffusion models learn personalized expressions? Global Representation Learning for Talking Portrait Generation

Abstract

In recent years, diffusion models have shown great potential in talking portrait generation. However, limitations still exist in personalized representation, particularly in appearance and motion, which we attribute to the current learning mechanisms. Regarding portrait appearance, the model tends to favor high-probability samples from the training data, leading to averaged results and weakened individual characteristics. In terms of motion, the model tends to produce locally plausible yet independent motion patterns, making it difficult to capture coordinated relationships in cross-identity portrait animation. To address these issues, we propose a talking portrait generation method based on global representation learning, which consists of two core components: alignment from group to individual (AG2I) and fusion of geometry and motion (FG&M). AG2I constructs the global appearance representations and performs group-to-individual alignment, thereby avoiding averaged appearance. FG&M constructs the global motion representations and leverages the fusion of geometric structures with multi-scale motion, enhancing motion coordination. Experimental results demonstrate that the proposed method effectively enhances the personalization and realism of the generated results compared to state-of-the-art methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.