FLARE: Fast Lightweight Avatars with Real-time Expression
Abstract
The deployment of interactive digital avatars, particularly for low-latency applications like LLM-driven interaction, is often hindered by a trade-off between rendering speed, visual quality, and generalization. While many methods achieve real-time frame rates, they can fall short of the high-throughput demands required for truly scalable and responsive systems. To address this, we present Fast Lightweight Avatars with Real-time Expression (FLARE), a one-shot 3D head avatar framework designed for high-speed synthesis. Our approach uses a decoupled Core-Shell Gaussian representation, with a dynamic FLAME-based Core and a static Shell for finer details. Crucially, we replace heavy per-frame neural deformation with a Hybrid Kinematic-Neural Deformation architecture. This relies on a one-time geometric binding based on barycentric coordinates as a gross kinematic base, paired with an lightweight Multi-Layer Perceptron (MLP) for high-frequency residual deformation and opacity control. This decoupled design effectively eliminates key computational bottlenecks in prior work. Consequently, our method achieves rendering speeds of over 600 FPS for motion reenactment and over 350 FPS for end-to-end speech-driven generation, measured at resolution on a single NVIDIA A100 GPU. Extensive experiments show FLARE achieves state-of-the-art inference speed while maintaining competitive reconstruction quality, offering a practical and scalable solution for deploying high-fidelity avatars in latency-sensitive applications. Code will be available upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.