Prototype Margins and Classifier Decisions: A Geometric Analysis of Deepfake Attribution Robustness
Abstract
Deepfake attribution models are vulnerable to image degradations, but whether prototype geometry reliably reflects classifier decisions remains unclear. We introduce Attribution Margin Collapse (AMC), a diagnostic framework that examines class-center displacement and compares prototype margins with classifier logit margins. Two complementary statistics, AMC-S and AMC-F, summarize the mean prototype margin and the dataset-wide magnitude of negative prototype margins, respectively. Experiments on FaceForensics++ across four attribution models document robustness degradation, while JPEG analysis reveals similar aggregate trends in prototype and classifier margins. A complementary cross-domain study of binary deepfake detection on Celeb-DF reveals substantial disagreement: samples can retain positive prototype margins despite incorrect classifier predictions. These findings highlight both the utility and the limitations of prototype geometry for robustness analysis. AMC provides a framework for examining when geometric diagnostics reflect classifier behavior and when they offer an incomplete account of failure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.