acceptodds
Under review as a conference paper at ICLR 2027

RuleSafe-VL: Benchmarking Rule-Conditioned Reasoning in Vision-Language Moderation

Abstract

Policy-grounded vision-language moderation can warrant different decisions for the same content depending on context. When evidence about outcome-changing conditions is missing, definitive judgments may rely on unsupported assumptions. How should policy application be formulated when conditions interact and the available evidence leaves multiple outcomes possible? We formalize this problem as a rule-conditioned reasoning task and introduce RuleSafe-VL. Derived from public platform policies through expert annotation and adjudication, its 93 atomic rules and 92 typed relations are linked to 2,166 expert-reviewed image–text cases across three high-risk domains. Four complementary tasks evaluate satisfied policy conditions, rule relations, decision states, and context-conditioned decisions. Across 10 frontier, open-source, and safety-oriented VLMs, the best rule-relation recovery and decision-state assessment scores reach 64.8 and 64.5 Macro-F1, respectively. Context-Pair Accuracy peaks at 35.4%, compared with an expert reference of 84.2%. Diagnostic analyses reveal modest gains in decision-state assessment from providing expert-annotated policy structure, alongside both missed policy-required decision changes and unnecessary changes when the correct outcome remains unchanged. These findings highlight challenges in determining when evidence warrants a moderation decision and how that decision should respond to policy-relevant context.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.