acceptodds
Under review as a conference paper at ICLR 2027

Real Images, Fake Verdicts: When Detectors Fail on Fake-Looking Photos

Abstract

Synthetic-image detectors are evaluated almost exclusively on benchmarks whose real images are typical photographs. We study the opposite tail: genuine photographs that look fake to people — extreme lighting, improbable geometry, coincidences, unusual post-processing. We introduce FLORIDA, 795 such curated real photographs, and evaluate ten detectors across six families after a shared control-benchmark sanity check. Seven pass; six nevertheless flag FLORIDA as fake at 1.8×–33× their control false-positive rate (all p < 10−4), up to 80% for the strongest detector. A provenance-matched control shows web provenance alone elevates FPR for five of seven detectors, while a genuine content effect survives matching for four of seven (up to 5.5×). The failures are also nearly independent across detectors (median φ = 0.04): 95.5% of FLORIDA fools at least one detector, yet none fools all seven. One deployed detector shows no significant elevation, and retraining a linear probe with 396 fake-looking reals as hard negatives cuts FPR from 55.9% to 1.0%: robustness is attainable and cheap. A nine-way taxonomy of why images look fake finds no safe category (union FPR 91–98%), with each detector sensitive to a different one; fake-looking reals are more atypical in frequency space than the fakes themselves. “Fake-looking” is not one shared cue detectors accidentally learn, but a broad region of real-image space where each detector fails differently — with direct consequences for deploying detectors as evidence of authenticity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.