DynRule: Benchmarking and Modeling Dynamic Rule Interactions for Universal Multimodal Content Moderation
Abstract
Content moderation models are typically optimized for limited rule patterns, making them less adaptable to the diverse and dynamically evolving hierarchical rule configurations encountered on real-world platforms. In practice, rules evolve with business requirements, interact under different authorization settings, and may be further complicated by untrusted pseudo-rules embedded in multimodal user content. Existing benchmarks typically study policy changes or instruction hierarchy in isolation, leaving the dynamic interactions among multiple rule levels underexplored. To better reflect real-world moderation, we introduce DynRule-Bench, which systematically varies hierarchical rule sources and their interactions as business requirements evolve. Controlled edits, source exchanges, and pseudo-rule insertions generate diverse rule interaction patterns across four scenarios, evaluating both adaptation and preservation. We further propose DynRule-Guard, a universal rule-interaction guidance framework. DynRule-Guard constructs a Rule Relation Graph to decompose complex hierarchical interactions into explicit local relations, aggregates them into a compact Global Interaction State in a shared Rule-Interaction Latent Space, and dynamically composes reusable patterns from an Interaction Pattern Bank to instantiate an interaction-specific Guidance Token. This design enables universal moderation across diverse and unseen rule configurations by inferring their interactions in a shared Rule-Interaction Latent Space and pattern composition, without rule-specific retraining. Extensive experiments on DynRule-Bench reveal the brittleness of existing moderation methods under dynamic hierarchical rules and demonstrate that DynRule-Guard consistently improves moderation performance, with further generalization to instruction-hierarchy evaluation on IHEval. Our code and data will be released publicly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.