Can AI Help Regulate AI? Benchmarking Multimodal Models as Advertising Regulators
Abstract
AI generation is scaling far faster than human regulation. Generative models enable regulated content to be produced at machine speed and near-zero marginal cost, with multimodal advertising generation providing a prominent example, while regulatory review remains constrained by scarce human attention. This growing asymmetry creates a **scalable regulation problem**: can AI help humans regulate AI-generated content at the scale at which it is produced? We introduce **RegulaBench**, a benchmark for evaluating multimodal models as advertising **pre-regulators**—scalable assistants that screen potentially unlawful content, organize legally relevant evidence, and provide grounded regulatory recommendations for subsequent human review. Built from real advertising enforcement cases, RegulaBench links generated advertising images to the factual findings and legal grounds recorded in corresponding administrative penalty decisions. It evaluates three progressively demanding capabilities: **violation detection**, whether models identify advertisements that trigger regulatory intervention; **evidence-to-norm grounding**, whether they connect case-specific evidence to the conditions governing applicable legal provisions; and **regulatory fidelity**, whether their selection and application of legal norms remain aligned with the structure of human regulatory decisions. Our experiments reveal a two-sided picture of AI-assisted regulation. On the one hand, across a substantial share of cases, model judgments are consistent with those of human regulators, suggesting that AI can help extend human regulatory capacity in screening potentially unlawful advertising at scale. On the other hand, this agreement becomes substantially weaker once we move beyond regulatory outcomes to examine how those outcomes are legally justified. In particular, across the provision groups examined, models show a directional tendency to rely more heavily on specific content prohibitions, whereas human enforcement decisions more often invoke broader truthfulness provisions. These findings reveal an important tension in scalable AI-assisted regulation: **AI can scale regulation without faithfully reproducing existing patterns of human regulation**. Such divergence raises consequential questions about how AI pre-regulators should be integrated into human regulatory systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.