From Typographic Distraction to Selective Token Intervention in CLIP
Abstract
Typographic attacks can redirect CLIP's prediction toward a rendered word even when the depicted object remains visible. This creates a selective intervention problem: suppress the distracting cue while retaining the image evidence needed for recognition. We study where this interference enters the visual representation through paired clean and attacked images. Restoring clean patch representations at positions overlapping the rendered text largely reverses the shift toward the rendered label, while inserting attacked representations at the same positions induces it. These findings motivate Patch Token Suppression (PTS), which learns a typographic score for each patch and moves representations with high scores toward a centroid estimated from patches assigned low scores in the same image. On synthetic ImageNet attacks, PTS improves object recovery while largely preserving clean accuracy. Evaluations on attacks in real images and ten additional datasets show consistent reductions in text confusion. Ablations identify spatial localization as the principal contributor to robustness, with smaller effects from the replacement target and interpolation rule.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.