Rationale Content Shapes What a Detector Learns: Controlled Supervision for Person-Centric AI-Image Forensics
Abstract
Detectors trained to generate rationales alongside authenticity decisions may learn different cues depending on the evidence they are taught to explain. This study focuses on consistency between gaze and attended targets in person-centric AI-generated images and analyzes how incorporating such relational gaze information into rationales changes detection performance and output behavior. We conduct experiments with Custom Gaze, 23,415 identity-preserving real/fake pairs that differ only in a generated eye band, and find that gaze-centered fine-tuning outperforms FakeVLM origin across three person-centric evaluation sets. The detector also maintains high balanced accuracy on eye edits from three held-out inpainters, and the largest point-estimate gain under person-count stratification appears in the stratum with two or more detected faces. It achieves the strongest aggregate metrics among the representative detectors compared, while paired-data training produces the same directional effect on the vision-only Effort detector. Finally, with the data, five-block structure, optimization, and checkpoint rule fixed, gaze-centered rationales outperform anatomy-centered and context-centered alternatives on COCOAI Person by 4.2–4.6 and 2.1–2.4 balanced-accuracy points, respectively. These results show that the evidence emphasized in a rationale is a substantive training factor that changes detector behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.