acceptodds
Under review as a conference paper at ICLR 2027

When Conversation is the Attack: Mining Transferable Contexts for Harmful Compliance and Over-Refusal

Abstract

Single-turn prompts are insufficient for evaluating the safety of agents that operate over extended conversations. We study automated conversation-history mining in which an attacker model iteratively constructs a conversation history preceding a fixed final request to discover contexts that induce either harmful compliance or an over-refusal. We evaluate a full 9×9 attacker–target panel using balanced sam- ples of 100 harmful and 100 benign prompts from WildGuardTest. For suscep- tibility, harmful-compliance discovery rates range from 1.2–14.5% while benign over-refusal discovery rates range from 0.7–21.3%. Claude 5 Sonnet and Muse Spark 1.0 are relatively resistant to harmful-compliance mining but susceptible to over-refusal mining, whereas Gemini 3.5 Flash shows the opposite. GPT-5.6 Terra and GPT-5.6 Sol are relatively robust to both objectives, while Grok 4.6 and 4.3 are susceptible to both. For attacker effectiveness, Grok 4.6 achieves the highest harmful-compliance discovery rate, whereas GPT-5.6 Terra achieves the highest over-refusal discovery rate. Self-attack provides no clear advantage. For trajec- tory transferability, successfully mined histories reproduce on their source mod- els in 52–57% of fresh generations but transfer to other models in only 14–17% under prefill-style last-turn replay. This large source-to-transfer gap shows that au- tomatically mined histories should not be assumed to be model-neutral benchmark items. Safety benchmarks based on automatic red-teaming methods can substan- tially favor or disadvantage models depending on which source models were used to supply histories.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.