TAVR4D: Temporal Audio-Visual Transformers for Reliable 4D Head Reconstruction
Abstract
Monocular 3D head reconstruction is accurate when visual evidence is reliable, but real videos often contain occlusions, illumination changes, motion blur, and other image degradations. Under these conditions, frame-wise estimators can interpret corrupted observations as motion, leading to large errors and jitter in pose and expression prediction. Temporal smoothing can reduce jitter but does not recover motion when the current observation is missing or unreliable. We introduce TAVR4D, a long-context multimodal model for robust 4D head reconstruction from degraded video. We formulate reconstruction as a temporal inference problem in which different motion components benefit from different complementary sources of information. The model uses long-range visual context to estimate head rotation, neck pose and translation when individual frames are unreliable. It also uses synchronized speech to help estimate facial expression and jaw motion. A learned visual reliability gate controls the balance between visual and audio features. We evaluate TAVR4D under 14 visual degradation conditions, including real-hand and full-face occlusions as well as dropped frames, which are held out as explicit image-space corruption classes during training. Across conditions shared by all baselines, TAVR4D remains close to its clean-input reconstruction accuracy, while remaining competitive on clean input. Our results show that combining long-range temporal inference with reliability-gated audio-visual fusion improves motion recovery and reconstruction robustness under the evaluated visual degradations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.