SCORE: Semantic-Conditioned Orthogonal Refinement for Generalizable AI-Generated Image Detection
Abstract
Generalizable forgery detection requires forensic evidence that remains reliable under shifts in content, acquisition conditions, and generative mechanisms. A central challenge is that the diagnostic meaning of many forensic cues depends on the visual context in which they occur: local texture, boundary, and frequency responses can vary with identity, pose, expression, scene structure, and image quality. This motivates a controlled use of semantic context when forming and refining detection decisions. We propose SCORE, a Semantic-Conditioned Orthogonal Refinement framework that organizes detection into a semantic-conditioned primary stage and a conservative refinement stage. SCORE first uses the current visual state to parameterize a gated residual and forms an independently supervised primary prediction. It then constructs a sample-conditioned low-dimensional basis and applies soft orthogonal attenuation to obtain a structured refinement view that preserves part of the basis-aligned information. A conditional refinement head operates on this detached representation, while uncertainty derived from the detached primary prediction controls the strength of a bounded logit correction. SCORE achieves an average video AUC of across five cross-dataset targets, an average video AUC of over eight unseen manipulations in DF40, and mAP and mAcc across 19 UniversalFakeDetect (UFD) test sources. These results support structured semantic conditioning and controlled refinement as an effective approach to forgery detection under diverse distribution shifts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.