acceptodds
Under review as a conference paper at ICLR 2027

Evidence Is Layered: Expert-Guided Queries for Generalizable Image Forgery Detection

Abstract

Pretrained vision transformers (ViTs) provide transferable representations for generalizable image forgery detection, yet how forensic evidence is distributed across their layers and how to retrieve it effectively remain poorly understood. We investigate these questions with LayerProbe, a controlled study of 14 frozen ViT backbones spanning standalone and MLLM vision encoders, using layer-wise linear probing and greedy layer fusion. The results reveal that the best readout depth varies across backbones, while selected layers with lower single-layer scores can still complement the strongest layer. Motivated by these findings, we propose ForenLens, a frozen-backbone detector that uses queries to retrieve complementary forensic evidence from a compact, backbone-specific layer set selected by LayerProbe. ForenLens uses a shared cross-attention reader to extract layer-specific responses before combining them through query-adaptive routing, preserving layer identity and learning where each query reads for each image. Complementing this adaptive layer routing, training-only forensic guidance shapes what expert queries retrieve by supervising their spatial attention. On the OpenMMSec out-of-distribution split and seven external benchmarks, ForenLens achieves state-of-the-art performance with 93.2% mean AP and 87.1% mean accuracy, while remaining robust to post-processing. Our code will be available upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.