Ranking Is Not Enough: Shared-Threshold Robustness in AI-Generated Image Detection
Abstract
AI-generated image detectors are often deployed on images that have been compressed, resized, blurred, transmitted, or re-digitized. These transformations can shift detector scores even when their within-condition ranking remains informative. As a result, strong conditionwise detection does not guarantee that one fixed orientation and threshold will work across all deployment conditions. We formalize this distinction in two stages. First, we study whether detector evidence itself survives processing. Our scalar response, , compares the image change caused by one application of a fixed reconstruction operator with the additional change caused by a second application. We derive a response-survival decomposition that tracks how clean real-versus-generated separation is retained or lost through processing distortion and reconstruction-processing interaction. Second, we characterize when all conditions admit one common decision rule. Exact threshold intersections determine shared-threshold feasibility, while total variation gives lower bounds on any decision restricted to the response. Empirically, we evaluate across two real-image sources, six generator families, and 32 processing conditions using independent calibration and confirmation. We then apply the same protocol to 58 publicly available detectors. The results show both regimes: some scores support a useful shared operating point, while many retain strong conditionwise ranking but fail under one fixed rule. For , the best shared threshold remains close to chance in its worst condition. These findings separate response survival from threshold compatibility and show why both are necessary for robust detector deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.