ModEval: Benchmarking Steering Modifiers for Social-Risk Evaluation in Text-to-Image Generation
Abstract
Text-to-image (T2I) models have been widely adopted for their impressive ability to generate high-quality images through textual prompts. During prompt composing, users often rely on prompt modifiers, i.e., auxiliary descriptions attached to the subject, to enhance visual richness and control artistic form. While modifiers provide a lightweight interface for steering generation, they may also alter model behavior in ways that lead to problematic outputs, raising broader social concerns. In this work, we propose ModEval, a unified evaluation framework for benchmarking the social impact of steering modifiers in T2I generation. We first harmonize modifier taxonomies into eight categories and construct the Prompt Modifier Dataset (PMD), which contains approximately 74,000 real-world prompts annotated with span-level modifier categories. Using PMD, we identify five steering modifier categories through dual-metric validation, combining image-level and semantic-level influence measurements. We then evaluate these steering categories across three social dimensions: demographic bias, deepfake detectability, and safety, on both open-source and black-box commercial T2I models. Our results show that all five steering categories can induce downstream social risks with different emphasis. Artist modifiers are primarily associated with demographic bias shifts; trending and medium modifiers substantially affect deepfake detectability; and atmosphere and movement exhibit broader multi-risk effects, with atmosphere most strongly linked to safety concerns.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.