CrossAnchor: Image-Anchored Text Optimization Exposes Blind Spots in Multi-line Defenses of Agentic Systems
Abstract
Red teaming of LLMs is still largely conducted by testing models in isolation. In reality, however, external guardrails typically stand as a first line of defense by screening and filtering harmful text inputs before they are passed to an LLM. In our work, we develop novel red-teaming attacks that test the multi-line defense as a whole. We build on the Greedy Coordinate Gradient (GCG) approach with two key differences. First, we swap the target-model log-likelihood loss of GCG for a cross-modal anchor, namely, the cosine distance to a non-text representation of the same intent. This is motivated by the asymmetry in a model's alignment across modalities. As the anchor we use a typographic render or a related image-only artifact, and we average the cosine distance across an ensemble of independently trained dual-encoders. This helps remove dependence of the optimization signal on any specific target model and produces transfer across architectures and modalities. The optimizer never sees the target model. Crucially, the anchor is needed only to construct the attack, not to apply it. Second, instead of appending a suffix, we search over a sparse, randomly selected subset of the prompt's own interior tokens, sparse enough to disrupt guardrail pattern-matching. The cross-modal loss recovers the steering signal that masking destroys. On prompts drawn from SALAD-Bench, evaluated across 8 sparsity levels and 6 frontier models, Qwen3Guard detection drops from 50.3% for unedited prompts to 1.2% for optimized prompts at 80% sparsity while optimized ASR at the same sparsity recovers from 4.3% for random-fill baselines to 14.3% to 68.3% across the six target models. Played as a sparsity cascade, the method succeeds end-to-end on 71.4% to 97.5% of the cohort across the six target models, including 71.4% and 97.5% on the two held-out models. This exposes a blind spot in typical evaluation scenarios and highlights the need for defenses that are robust to techniques with non-textual optimization signals. The optimized prompts and end-to-end success envelope form an adversarial corpus and residual-risk benchmark to harden input-side guardrails of agentic systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.