RuleSafe-VL: Benchmarking Rule-Conditioned Reasoning in Vision-Language Moderation
Abstract
Policy-grounded vision-language moderation can warrant different decisions for the same content depending on context. When evidence about outcome-changing conditions is missing, definitive judgments may rely on unsupported assumptions. How should policy application be formulated when conditions interact and the available evidence leaves multiple outcomes possible? We formalize this problem as a rule-conditioned reasoning task and introduce RuleSafe-VL. Derived from public platform policies through expert annotation and adjudication, its 93 atomic rules and 92 typed relations are linked to 2,166 expert-reviewed image–text cases across three high-risk domains. Four complementary tasks evaluate satisfied policy conditions, rule relations, decision states, and context-conditioned decisions. Across 10 frontier, open-source, and safety-oriented VLMs, the best rule-relation recovery and decision-state assessment scores reach 64.8 and 64.5 Macro-F1, respectively. Context-Pair Accuracy peaks at 35.4%, compared with an expert reference of 84.2%. Diagnostic analyses reveal modest gains in decision-state assessment from providing expert-annotated policy structure, alongside both missed policy-required decision changes and unnecessary changes when the correct outcome remains unchanged. These findings highlight challenges in determining when evidence warrants a moderation decision and how that decision should respond to policy-relevant context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.