acceptodds
Under review as a conference paper at ICLR 2027

UniPortrait4D: Unifying 4D Portrait Reshooting and Animation

Abstract

Portrait animation and reshooting (novel-view synthesis) have been extensively studied for video generation, yet existing approaches typically treat them as separate problems: expression-driven animation is often restricted to a fixed viewpoint, while reshooting relies on limited multi-view training data with insufficient identity diversity. We present UniPortrait4D, a unified video DiT that supports both synchronized multi-view reshooting from a monocular portrait video and 4D portrait animation. To unify 4D portrait reshooting and animation, we leverage FLAME-based rendering as a unified camera representation for multi-view generation. For reshooting, we additionally incorporate point-cloud renderings to improve multi-view consistency. For animation, we employ a high-capacity latent representation for facial expressions and adversarially train it with gradient reversal to reduce driver-identity predictability while preserving fine-grained facial motion. Benefiting from the 4D consistency of our video DiT, the generated multi-view videos can be reconstructed into a 4DGS representation, enabling real-time free-viewpoint rendering. Extensive experiments demonstrate that our method achieves strong reference-identity preservation in cross-identity animation and high multi-view consistency in reshooting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.