acceptodds
Under review as a conference paper at ICLR 2027

AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures

Abstract

Modeling pose-dependent surface dynamics is essential for realistic full-body avatars. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles. In contrast, Person-specific methods capture such dynamics but require costly multi-view acquisition for each identity. Generalizable dynamic methods either produce limited surface dynamics or sacrifice multi-view consistency when introducing them. We present AvatarDynamizer, a generative method that transforms an off-the-shelf static avatar into a controllable 4D avatar with realistic surface dynamics and multi-view-consistent rendering. Our key insight is to represent pose-dependent surface dynamics as dynamic textures, bridging pretrained video diffusion models and 3D Gaussian splatting. Specifically, our Dynamic Texture Generator produces fine-grained, motion-dependent textures, which are decoded by a Generalized Gaussian Decoder into 3D Gaussians for inherently multi-view-consistent rendering. To bridge the gap between identity-rich but motion-limited datasets and motion-rich datasets with few subjects, we collect a high-quality multi-view dataset with dense camera coverage and diverse identities and motions. Experiments across four datasets show that our method outperforms competing generalizable dynamic methods on generative metrics and produces more faithful high-frequency surface dynamics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.