GROVE: Governed Rule Operationalization through Verifiable Agentic Evolution for Multimodal Moderation
Abstract
How can a multimodal moderator adapt to an authorized policy revision without retraining its vision–language model or silently changing unrelated rules? We introduce GROVE (Governed Rule Operationalization through Verifiable Agentic Evolution), a two-stage framework that compiles policy clauses, exceptions, and scopes into a versioned schema of atomic questions. Its release protocol specifies source-bound agent edits and independent checks of policy fidelity and protected behavior. At inference, five-choice answer-token distributions supply rule-linked evidence to a lightweight decision adapter. We evaluate this evidence path on five public image–text benchmarks with group-isolated predictions and fixed-readout controls. Adding atomic evidence to a fixed XGBoost head improves macro-F1 by 3.93 points on MAMI, ROC-AUC by 1.56 points on Hateful Memes, and reject F1 by 8.69 points on HARM-C relative to the same head using only the base score. A policy-free representation head serves as a demanding control. The measured contribution is an informative, traceable evidence channel; sequential policy-update evaluation is specified separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.