Balanced Perturbation and Subspace Erasure Do Not Identify What an Activation Monitor Reads
Abstract
Activation monitors score behaviours such as deception from language-model hidden states, but held-out accuracy cannot show whether a monitor reads the behaviour itself or a correlated feature. We analyse two intervention diagnostics for this distinction. Under a balanced perturbation with conditional mean zero, positive and negative first-order score responses cancel in expectation. Its mean score change therefore reflects local curvature rather than target reliance. For linear monitors under elliptical perturbations, variance, mean absolute change, and flip rate are also non-identifying. Matched erasure compares the score loss after removing a fitted class-separating direction with the loss from an equal-size control edit. A monitor that reads a correlated attribute can share the same alignment with this direction as a target reader. In panels with sealed construction labels, matched erasure ranks readers above confound-riders, but no threshold separates them, including after site-level normalization. A margin calibrated on one corpus passes of riders in a corpus-matched follow-up; all fail when their attribute is reversed while the behaviour is held fixed. Two released deception probes pass in only of benchmark cells. Reversing a named candidate attribute classifies of monitors correctly on held-out records. We recommend reporting this test with its full operating curve rather than treating one intervention result as evidence of target reliance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.