Jailbreaking LLM-based Agents via Multimodal Skills
Abstract
Agent skills have recently emerged as a powerful mechanism for extending the capabilities of LLM-based agents, but they also introduce a new jailbreak surface. Existing work has shown that harmful textual skills can weaken agent refusal, while multimodal resources can evade skill scanners when used to conceal additional malicious instructions. We identify a complementary vulnerability: agent-skill safety can change when the same harmful source payload is delivered through a different multimodal representation. We introduce ModalityShift, a content-preserving multimodal skill attack that transforms an existing harmful textual skill into an ordered sequence of images without generating, deleting, reordering, or rewriting its source payload. The image-based payload is delivered through the instruction required by the multimodal agent interface. ModalityShift searches for representations that suppress refusal-related signals in an accessible proxy model and transfers the resulting payloads to black-box target agents without querying the target models or skill scanners during optimization. We evaluate the resulting deliveries at both runtime and pre-deployment boundaries. Across four target agents, ModalityShift produces model-dependent reductions in refusal on its own and substantially amplifies existing multimodal jailbreaks on several models. Meanwhile, image-based delivery sharply reduces unsafe decisions from representative skill scanners. These results reveal a representation-dependent safety gap in current agent-skill pipelines and motivate defenses that reason consistently over skill semantics across delivery modalities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.