NOVAD: Variable Native Observation for AI-Generated Image Detection
Abstract
AI-generated images appear at widely varying resolutions and aspect ratios, yet fine-grained forensic cues are better preserved at native scale. Resizing large images can suppress these cues, motivating detectors to encode large images as local views at native scale. However, exhaustively processing all native regions makes inference cost grow with image area. This raises a natural question: can a detector preserve access to native detail without requiring full spatial coverage? We introduce **NOVAD** (**N**ative **O**bservation with **V**ariable **A**llocation for **D**etection), a detector that learns to operate on subsets of native local views together with a global view. During training, NOVAD is exposed to different numbers and spatial configurations of local views, teaching a shared detector to combine global features with local evidence from changing view subsets. These subsets reuse features encoded once per view, enabling multi-set supervision without repeated view encoding. Geometry-aware fusion combines local detail, whole-image context, and spatial information, allowing one model to handle changing view sets, resolutions, and aspect ratios. At inference, attribution from the global-only prediction guides local tile selection before local encoding, while the inference budget determines how many tiles are processed. The same model operates across inference budgets without retraining. Experiments show that NOVAD retains strong detection performance with only a few native local views. Trained on OpenFake, it achieves **98.44%** mean accuracy across Chameleon, CommunityAI, SocialRF, and WildRF, and **98.05%** unweighted mean accuracy across the eight resolution strata of HiRes-50K.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.