One Feature Label, Many Recipes: Auditing Jailbreak Feature Claims Across Prompt Versions and Environments
Abstract
Large language models (LLMs) remain vulnerable to jailbreak attacks that elicit harmful responses. Prior work uses interpretable prompt features to explain jailbreak behavior and guide attack and defense design. However, one feature label can correspond to many concrete prompt edits (recipes), making it unclear when evidence from individual recipes supports a broader feature claim. We propose EdgeAudit, a framework that jointly specifies how recipe effects may be aggregated and how the resulting evidence may be interpreted. The key idea is to represent a recipe family by the set of effects induced by its permitted selection policies, and to tie the corresponding confidence bounds to the evidence each interpretation level requires. This formulation distinguishes support for an average over recipes from support for every allowed recipe, and claims about an assigned edit from claims about the intended feature or an edge in a learned feature graph. It thus preserves supported local findings while identifying the additional evidence required for stronger interpretations or extensions to other environments. In paired audits on Mistral at a 200-token output budget, requesting five numbered steps instead of a continuous response raises the HarmBench-judged harmful-compliance rate by 0.462 in plain context but by only 0.115 under jailbreak scaffolds from JailbreakHub. On held-out Qwen behaviors, changing the requested count from one to five has an effect 0.155 larger in a checklist than in a JSON array. Effects attributed to one feature label can therefore depend on context and format.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.