acceptodds
Under review as a conference paper at ICLR 2027

Occlusion-Aware Dual-View Masked Autoencoding for Facial Representation Learning

Abstract

Visual intelligence is increasingly shifting from world-centric perception toward human-centric understanding. However, the visual evidence required to interpret human behavior and affective states remains susceptible to occlusion-induced corruption. Existing masked pretraining paradigms leave this setting underexplored, as their objectives can be satisfied by recovering missing facial content from visible context without requiring the encoder to distinguish facial evidence from occluder-induced appearance. When corrupted regions remain visible, the pretraining objective does not explicitly constrain the encoder against treating occluder-specific patterns as valid facial evidence, allowing them to persist alongside behavior-relevant facial information in the learned representation. We propose a dual-view masked autoencoding framework that learns from paired clean and synthetically occluded facial views. In the occluded branch, corrupted patches remain visible and are guided to reconstruct their corresponding clean facial content rather than the occluder appearance, directly linking corrupted observations to the facial information they obscure. This encourages the encoder to preserve facial semantics while reducing reliance on occluder-specific cues. Experiments on facial action unit detection and facial expression recognition demonstrate the effectiveness of the learned representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.