SPATIAL PROBING FOR LOCALIZED AI-EDITED IMAGE DETECTION
Abstract
Frozen vision-foundation-model (VFM) features with a linear probe detect fully generated images accurately, but fail on localized editing: real photographs in which an AI generated, inserted, or re-rendered part of the content. We show most of this failure is an aggregation artifact: the pooled feature is an attention weighted average over spatial content, so on an edited image it aligns more with the unedited majority (cosine 0.233 to the all-patch mean) than with the edited patches (0.100). Moving the probe to the spatial level, as multi-scale windowed scoring or per-patch token probing, recovers much of the signal with the backbone frozen and no added training. With the released MetaCLIP2 probe (we reproduce its published benchmark tables to within 0.001 average accuracy), windowed scoring reaches 0.937 AUC on FLUX-Kontext instruction-edited images, 0.850 AUC cross-domain on BR-Gen, and stays above 0.89 AUC under JPEG and blur; JPEG re-compression even improves detection. Across the AIGI-Now matrix, windowed wins all 9/9 semantic splits (sign-test p=0.002); pooled wins 8/9 pristine pixel splits. Across three benchmarks, max aggregation captures the strongest local evidence and mean aggregation the global statistics; the operative choice is spatial isolation, not feature granularity, since whole-image patch statistics dilute a localized edit while windowed cropping isolates it. The windowed probe also estimates the manipulation’s spatial scale via its argmax crop. The paradigm’s boundary remains honest: patch-token probes cannot localize small edits (IoU near 0), whose signal lives in native-resolution noise statistics that downsampling destroys.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.