VideoFace: Learning a Discrete Motion Space for Sparse-to-Dense Facial Motion Reconstruction
Abstract
Reconstructing facial motion from sparse visual keyframes requires visual estimation under appearance changes and temporal completion between observations. Independently estimated states and subsequent interpolation do not explicitly model how an expression develops, including its relative intensity, regional coordination, and timing. We present VideoFace, a framework that combines visual conditioning from a pretrained vision-language model with a learned discrete motion prior. Motivated by the potential of semantic representations to support robustness across appearances, the model uses keyframe images, their temporal indices, and motion history to predict a complete motion-token sequence of a specified length. A dual-branch tokenizer connects ARKit and FLAME through shared finite scalar quantization, assigning one token per frame. Self- and cross-reconstruction preserve motion content, while window-level code matching and frame-level alignment encourage correspondence between the two representations. After tokenizer training, we freeze its parameters and fine-tune the vision-language model for autoregressive token prediction; either decoder then produces the corresponding continuous trajectory. We also construct MAMC-Face, comprising 10 hours of facial motion at 30 Hz rendered across 65 avatars under moving camera views. Its paired visual–motion supervision supports learning complete motion trajectories from sparse visual observations across varied appearances and viewpoints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.