acceptodds
Under review as a conference paper at ICLR 2027

Prototype Margins and Classifier Decisions: A Geometric Analysis of Deepfake Attribution Robustness

Abstract

Deepfake attribution models are vulnerable to image degradations, but whether prototype geometry reliably reflects classifier decisions remains unclear. We introduce Attribution Margin Collapse (AMC), a diagnostic framework that examines class-center displacement and compares prototype margins with classifier logit margins. Two complementary statistics, AMC-S and AMC-F, summarize the mean prototype margin and the dataset-wide magnitude of negative prototype margins, respectively. Experiments on FaceForensics++ across four attribution models document robustness degradation, while JPEG analysis reveals similar aggregate trends in prototype and classifier margins. A complementary cross-domain study of binary deepfake detection on Celeb-DF reveals substantial disagreement: samples can retain positive prototype margins despite incorrect classifier predictions. These findings highlight both the utility and the limitations of prototype geometry for robustness analysis. AMC provides a framework for examining when geometric diagnostics reflect classifier behavior and when they offer an incomplete account of failure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.