Multimodal DeepFake Detection via Domain-Incremental Multi-Adapter Learning
Abstract
DeepFake detectors are typically trained in a static setting and struggle to generalize to emerging manipulation techniques. As generative methods evolve rapidly, the ability to incrementally extend models to new domains without full retraining becomes critical. We introduce DEFEND, a framework for domain-incremental multimodal DeepFake detection with parameter-efficient updates. Unlike prior work that focuses on dataset-level shift, we consider a fine-grained incremental setting, where new manipulation types arrive sequentially. To this end, we construct a chronological stream of 23 domains, spanning from early face-swapping methods to recent multimodal forgeries. DEFEND builds on a frozen pretrained audio-visual backbone and leverages Low-Rank Adaptation (LoRA) to enable efficient domain-specific updates while mitigating catastrophic forgetting. Importantly, we introduce a dynamic adapter expansion guided by a learned key-query routing strategy. This mechanism allows the model to reuse existing adapters for similar manipulations while expanding its capacity when encountering novel ones. We evaluate DEFEND on three benchmarks, KoDF, DeepSpeak and MAVOS-DD. On BRAVEn, DEFEND reaches 89.7% average AUC under automatic routing, 8.5 points above the strongest continual learning baseline, with only 11 adapters for 23 domains; on CLIP+CLAP, it is on par with the best continual learning baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.