Mitigating Many-shot Jailbreak Attacks with One Demonstration
Abstract
Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by conditioning them on many harmful question-answer demonstrations. We study whether this contextual influence can be countered during inference while preserving useful responses. We introduce Safend, an inference-time defense that restores refusal behavior with a single fixed safety demonstration. We characterize MSJ and Safend as competing contextual response patterns. Harmful demonstrations repeatedly associate unsafe requests with compliant responses, while Safend inserts a fixed safety-aligned refusal demonstration after the harmful examples and immediately before the target query. This view treats harmful compliance and refusal as competing response patterns whose influence depends on contextual aggregation rather than example count alone. Safend requires no parameter updates, request-specific optimization, or white-box access, and reuses the same demonstration across inputs. Experiments on open-weight and API-based LLMs show that Safend reduces the average attack success rate of context-based jailbreaks from 76.7% to 2.3% on Llama-3.1-8B-Instruct and from 46.6% to 2.3% on GPT-4o. These reductions are achieved while preserving general capabilities, long-context performance, and benign response behavior. These results establish a simple and practical approach for recovering safety from adversarial contextual adaptation at inference time. Code is available at https://anonymous.4open.science/r/Safend.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.