acceptodds
Under review as a conference paper at ICLR 2027

Red-Teaming Procedural and Persuasive Text Rendering in Text-to-Image Models

Abstract

Recent text-to-image (T2I) models can now render accurate and legible text within generated images. This capability allows visually benign images to convey dangerous procedures or deceptive messages through rendered text. However, existing red-teaming methods treat such text as a single risk source, conflating actionable guidance with persuasive messaging. Moreover, their black-box prompt-optimization approach fails to adequately capture the intricate relationship between prompt modifications and the resulting output states. To disentangle procedural and persuasive harms, we construct ToxiGlyph, a red-teaming dataset that pairs procedural and persuasive instances sharing the same harmful intent, spanning 10 risk categories and 31 subcategories. We further propose GlyphPilot, which formulates black-box prompt revision as a vision-guided multi-directional search. Rather than refining a single candidate along a fixed path, GlyphPilot generates and evaluates diverse revision strategies in parallel and uses visual feedback to select the most promising direction at each iteration. Across five commercial T2I models (e.g., Nano Banana 2 and GPT Image 2), GlyphPilot achieves an 87.85% ASR, surpassing the strongest baseline by over 40 points, with persuasive text consistently more harmful than procedural text. These results establish text rendering as a critical yet underexplored attack surface in T2I models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.