SAE Interventions Can Be Unreliable: Post-Intervention Recovery of Suppressed Behavior
Abstract
Sparse autoencoders (SAEs) decompose residual-stream activations into interpretable features and increasingly support control mechanisms that monitor or clamp features associated with unwanted behavior. When a clamp suppresses its target behavior, the intervention is commonly treated as successful. This evaluation leaves a stronger question unresolved: has the behavior become inaccessible, or has one visible route merely been blocked? We introduce post-intervention recovery, a constrained residual-space audit of whether an SAE intervention forms a complete behavioral bottleneck. Starting from a successfully defended state, the audit searches for bounded residual perturbations that restore the suppressed behavior while the exact clamp remains active throughout optimization and generation. Encoder-orthogonal and feature-map Jacobian projections limit direct movement of defended readouts. Across Targeted Probe Perturbation (TPP), unlearning, Indirect Object Identification (IOI), and refusal steering, the audit repeatedly recovers behaviors that successful feature interventions had suppressed. In the safety-critical AdvBench setting, 20.7% of the tested valid flips become HarmBench-classified harmful under the active clamp. Frozen recovery directions transfer to held-out prompts and substantially outperform equal-norm random controls, while matched additive-steering defenses withstand the same search. Together, these results provide counterexamples to treating selected SAE features as complete behavioral bottlenecks under the evaluated perturbation classes. Behavioral completeness is therefore a distinct validation target for SAE-based control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.