Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs
Abstract
Jailbreak attacks on safety-aligned LLMs are increasingly effective, yet attack success alone does not reveal whether compliance is driven by preserved task intent, suppressed safety signals, or attack-specific artifacts. We use Pearl’s front-door adjustment to motivate separating task-relevant semantics from safety-related variation. We proposeCausal Front-Door Adjustment Attack (CFA), a white-box framework that operationalizes this perspective in LLM latent space using Sparse Autoencoders, contrastive datasets, and weight orthogonalization. After offline calibration, CFA enables efficient generation without per-prompt iterative optimization, performs strongly on HarmBench across four open-source LLMs, and provides source-side signals for same-family black-box prompt transfer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.