acceptodds
Under review as a conference paper at ICLR 2027

OR-Talk: Occlusion-Robust Lip Synchronization via Selective Appearance Transport

Abstract

Audio-driven lip synchronization modifies mouth motion according to a target speech signal while preserving identity and head motion. Under facial occlusion, existing methods often distort or remove foreground occluders, since preserving occluder appearance conflicts with audio‑driven control of mouth shape. We propose OR-Talk, which formulates occluder recovery as an appearance transport problem. Based on this formulation, we introduce Occlusion Adaptive Transport (OAT) to selectively transfer the appearance cues needed for occluder reconstruction from the identity features to the decoder. To avoid leaking mouth information during this process, we construct paired identity conditions with matched head poses but different mouth shapes. This design preserves occluder appearance while keeping the mouth shape controlled by the input audio. Experiments show competitive performance on standard benchmarks and substantially improved robustness to facial occlusion. Under 50–100% mouth occlusion, OR-Talk improves PSNR by 6.65 dB and reduces OCC-E by 51.6% compared with X-Dub.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.