SemGS-Avatar: Semantic-Aligned Disentangled Generation of 4D Human Avatars from a Single Image
Abstract
Disentangled generation of 4D human avatars from a single image has achieved remarkable progress in recent years. However, existing methods suffer from two challenges: (1) Spatially, they struggle to accurately assemble separate body parts due to unclear rules on how parts fit together, leading to overlapping parts and broken boundaries. (2) Temporally, they fail to maintain smooth motion because estimated poses jump from frame to frame, causing the reconstructed avatar to shake and drift. To address these challenges, we propose SemGS-Avatar, a framework for high-fidelity 4D human avatar generation through semantic-aligned disentangled Gaussian splatting. Specifically, SemGS-Avatar consists of two core modules: Semantic-Aligned Gaussian Composer and Temporal Consistency Pose Optimization. The former explicitly aligns the semantic attributes of 3D Gaussians with the structured somatic manifold of the SMPL-X model to ensure region-consistent part binding and precise structural arrangement, enabling accurate body-part assembly and boundary refinement. The latter integrates a motion interpolation network with sliding-window kinematic smoothing to stabilize dynamic poses, ensuring smooth and coherent animation. Benefiting from the above designs, our method not only produces high-fidelity 4D human representations with sharper textures and cleaner part boundaries, but also reduces pose jitter and improves temporal stability. Extensive experiments demonstrate that our SemGS-Avatar outperforms existing methods in both visual fidelity and video-driven animation stability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.