AvatarBoost: A Universal Real-Time Video Enhancement Model for 4D Avatarization
Abstract
High-quality 4D portrait animation requires photorealistic appearance, consistent novel-view synthesis, and faithful motion control. While explicit 3D representations, such as 3D Gaussians, naturally support efficient rendering and reliable geometric consistency, feed-forward avatar models often produce videos with blurred textures and boundary artifacts. Conversely, video diffusion models can synthesize rich visual details, yet they lack explicit 3D constraints and suffer from costly multi-step sampling. To bridge this gap, we present AvatarBoost, a model-agnostic video enhancement framework that boosts the visual quality of existing 4D avatar models. Operating directly on rendered video outputs, AvatarBoost enhances their appearance without altering the underlying avatar representations or animation controls. In this framework, the rendered videos naturally provide viewpoint and motion guidance, which we complement with multiple reference images and their corresponding facial landmark maps to supply rich identity and facial-structure cues. Conditioned on these joint inputs, a video diffusion transformer learns to suppress rendering artifacts and faithfully restore fine details. To ensure practical efficiency, we further distill the enhancement model into a two-step generator using distribution matching distillation (DMD). Extensive experiments on VFHQ and NeRSemble demonstrate that a single AvatarBoost model improves the perceptual quality of both LAM and UIKA without backend-specific fine-tuning, reducing LPIPS by up to 42.1% and FID by up to 64.4% in self-reenactment. AvatarBoost achieves an inference throughput of 20 FPS on a single NVIDIA H800 GPU. An anonymous project page is available at https://anonymousauthor021.github.io/AvatarBoostProject/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.