Calibrate Elsewhere, Verify Here:Disjoint In-Scene Calibration for Multimodal Face Anti-Spoofing
Abstract
Multimodal face presentation attack detection must distinguish presentation effects from ordinary changes in illumination and sensing. Target-only inference can marginalize an unknown acquisition state or estimate it from the face; the former adds no sample-specific information about the current capture, while the latter can suppress nuisance-aligned attack evidence. We introduce SceneCal, which learns type-conditioned responses of heterogeneous non-target regions to a shared acquisition state. Controls determine a posterior that is fixed before the face is scored, with uncertainty propagated to partially observed target modalities. We analyze plug-in attenuation, target-only marginalization, and disjoint control conditioning in one Gaussian model. On public WMCA RGB–Depth leave-one-attack-out, SceneCal reaches 2.99% ACER, versus 3.37% for CTNet and 3.54% for MMDG++; on HQ-WMCA it reaches 9.21%, versus 9.88% and 10.23%. On HQ-MMFAS, CTNet is stronger on known instruments (2.06% vs. 3.08%), while SceneCal reaches 4.72% on held-out instruments, compared with 5.08% for CTNet and 5.45% for DADM.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.