CAEIR: Cross-Architecture Experts with Interaction-aware Routing for Deepfake Detection
Abstract
Generalizable deepfake detection requires identifying weak and heterogeneous forensic cues that vary across manipulation methods and data distributions. Existing mixture-of-experts detectors, however, typically specialize experts within a single architectural family, while expert-selection decisions are often based solely on the input representation, without explicitly comparing the forensic evidence produced by heterogeneous experts. Our key insight is that routing should depend not only on the input, but also on how heterogeneous experts respond to each sample. We introduce CAEIR, a cross-architecture framework that places convolutional, Transformer, and state-space experts in a unified competitive pool, summarizes their responses as evidence nodes, and models interactions among these nodes to produce expert-specific routing corrections that refine an input-conditioned Top-1 ranking. A separate bounded modulator scales the selected response without altering the routing decision. All candidate responses are computed to support interaction-aware expert selection, while only the selected response is injected into the detection residual. Trained on FF++ real videos with online synthetic fake supervision, CAEIR achieves video-level AUCs of 96.61%, 95.71%, 82.19%, and 89.76% on Celeb-DFv2, DFD, DFDC, and FFIW, respectively. Controlled ablations further support the benefits of cross-architecture diversity and response-conditioned routing beyond full-candidate evaluation alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.