TIPGuard: A Unified Suite for T2I Prompt Guardrails with Customizable Safety Policies
Abstract
Guardrail models are a cornerstone of trust for deploying AI at scale. In text-to-image (T2I) scenarios, they face a fundamental challenge: mapping abstract safety norms into accurate, fine-grained judgments of open-ended visual content while adapting to evolving safety policies. However, existing approaches often rely on costly data collection and human annotation, and their safeguards remain tied to predefined safety policies. To address these challenges, we introduce TipGuard, a unified framework spanning data synthesis, model training, and fine-grained evaluation. At its core is an almost fully automated synthesis pipeline that produces semantically rich and diverse safety supervision for realistic T2I requests with minimal human involvement. Using this pipeline, we construct, to the best of our knowledge, the largest T2I guardrail dataset, with 90% agreement between synthetic annotations and human judgments. We further train a series of guardrail models tailored to different deployment scenarios. These models achieve state-of-the-art performance across multiple safety benchmarks and demonstrate surprisingly strong generalization to previously unseen safety policies. We further introduce TipGuardBench, a benchmark for multi-category, fine-grained safety judgments, and use it to systematically assess the capabilities and limitations of existing guardrail models. Together, these findings demonstrate the feasibility of building policy-customizable T2I guardrails with largely automated supervision. By unifying data synthesis, model training, and evaluation around explicit safety policies, TipGuard provides a foundation for developing guardrails whose safety criteria can be revised as deployment requirements evolve.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.