When Attribution Graphs Fail to Predict LLM Jailbreaks: Distinguishing Structural Signals, Causal Effects, and Behavioral Safety
Abstract
Internal attribution graphs offer a promising way to examine how large language models process adversarial prompts, but their ability to predict, explain, and mitigate jailbreak behavior remains unclear. We investigate these three objectives separately using full-depth, transcoder-based MLP-feature attribution graphs for Llama-2-7B-chat. Across 50 topically adjacent clean–adversarial prompt pairs, we align graphs by exact feature identity and evaluate four structural measures: graph deviation, feature suppression, feature emergence, and path rerouting. None reliably distinguishes successful from failed jailbreaks under the evaluated setting, and the binary outcomes themselves are sensitive to decoding: only two of five successes observed with sampled decoding persist under greedy decoding. We next rank features by gradient-based sensitivity and screen them through targeted interventions. Ablating the selected features produces substantially larger changes in target-token loss than matched random-feature interventions, demonstrating that locally influential internal components can be identified even when aggregate graph statistics are not predictive. However, follow-up evaluation of complete generations shows that satisfying the loss-based intervention criterion does not reliably restore refusal behavior. These results reveal a three-way dissociation between structural prediction, causal sensitivity, and behavioral mitigation. Internal attribution graphs, as operationalized here, can localize computation affecting a target prediction but do not yet provide reliable jailbreak detectors or behaviorally validated defenses. Our findings highlight the need for decoding-stable outcome definitions, behavior-level intervention evaluation, and independently validated structural signatures in mechanistic studies of language-model safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.