acceptodds
Under review as a conference paper at ICLR 2027

Moderating Covertly Sensitive Images Across Policies: Policy-Reconfigurable Benchmark and Zero-Shot Method

Abstract

Harmful image moderation is an important task in content safety, complicated by ambiguity in what constitutes harmful content and variation across moderation policies. Existing datasets and methods often provide limited support for these challenges. To fill this gap, we introduce the Covert Sensitive Image Dataset (CSID), comprising 2,510 images collected from diverse sources and spanning semantically complex violation themes, with labels configurable under different moderation policies. By formulating theme-specific objective questions, CSID enables the reuse of a single set of annotations across policies. We further propose VLGuard, a training-free, evidence-guided framework for image moderation. VLGuard decomposes each image into structured semantic units and retrieves relevant sensitive concepts to inform moderation decisions by a vision-language model (VLM). Evaluation against 12 image moderation methods from industry and academia demonstrates that VLGuard achieves an average F1 score of 74.8% on CSID, outperforming the strongest evaluated academic baseline and the evaluated closed-source VLM by 16.7 and 3.7, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.