Beyond Homogeneous Gaussian Modeling: Is Cross-Region Geometry Representation the Key to Talking Portrait Generation?
Abstract
3D Gaussian Splatting (3DGS) demonstrates strong representation capability and great potential for portrait generation. However, generated portraits still suffer from significant geometric degradation. Specifically, due to the lack of global constraints, the geometric boundary between motion regions and the static background is difficult to maintain, making the boundary prone to deformation or distortion. Furthermore, structural characteristics in the portrait can lead to motion differences across regions, resulting in structural misalignment or discontinuities. We attribute these issues to the homogeneous representation adopted by existing 3DGS-based methods, which treats regions with different motion characteristics in a unified manner and fails to capture their distinct geometric properties. To address this problem, we propose a cross-region geometric representation framework that aims to characterize the geometric relationships among different motion regions for dynamic portrait generation. The proposed framework consists of two core strategies: global motion depth representation (GMDR) and local motion multi-scale representation (LMMR). GMDR constructs a depth representation of global motion regions to describe the spatial hierarchy between the dynamic foreground and static background, thereby alleviating geometric distortions. LMMR exploits multi-scale representations of local motion regions to capture structural variations in portraits, promoting motion coordination and structural integrity across regions in continuous sequences. Experimental results demonstrate that the proposed method significantly improves the geometric fidelity of portraits in continuous dynamic sequences and outperforms existing state-of-the-art approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.