Free360Head: Continuous-View Full-Head Generation via 3D-Aware Guidance
Abstract
Generating a continuous free-viewpoint video of a complete human head from a single portrait requires plausible completion of unseen regions, identity preservation, precise trajectory control, and cross-view consistency. Existing 3D-based approaches provide consistent rendering but are limited by the quality of single-image reconstruction, whereas generative approaches produce realistic details but may drift in geometry and identity under large viewpoint changes. We present , to the best of our knowledge the first reference-independent 3D-aware trajectory control framework, combining the strengths of both routes. Our method explicitly disentangles identity from trajectory: an arbitrary-view reference portrait specifies identity without defining any output viewpoint, while an identity-agnostic depth sequence independently controls the target trajectory. Compared with camera-pose conditioning, the depth sequence provides dense, frame-aligned cues for projected silhouette, visible surface, and depth ordering, enabling fine-grained geometric control. Built on the pretrained Open-Sora video diffusion Transformer, uses a reference-prefixed latent injection scheme and an auxiliary reference reconstruction objective to propagate identity throughout depth-guided video generation. Experiments demonstrate improved visual fidelity, identity preservation, cross-view consistency, and trajectory control over existing approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.