CLEAR: Calibrated Linear Erasure Auditing of Representations
Abstract
What does a performance drop after representation erasure reveal about a model's use of the targeted information? Erasure aims to remove target information from internal activations, but can also disrupt other task-relevant features. The resulting performance drop can therefore reflect both effects. Interpreting this drop requires checking how much target information remains, how sensitively residual tests can detect it, and how intervention choices affect behaviour. We introduce CLEAR, a framework for calibrated linear erasure auditing of representations. CLEAR interprets behavioural change together with the remaining linear target association, the sensitivity of the residual test, and the response to different intervention strengths and replacement values. Calibration makes the resolution of each residual verdict explicit, connecting the validity of an internal edit to the behavioural evidence it provides. Experiments demonstrate how this joint assessment reveals ambiguities hidden by effect magnitude alone. Synthetic controls show that error from fitting the eraser can produce behavioural effects even at zero target dependence. In the SSR planner, full ego state erasure raises 3 s average displacement error to 7.3 times its baseline value while failing residual checks at every audited layer. Error also grows beyond the intervention strength that minimizes target association and varies with replacement values. CLEAR makes explicit how residual information, detection sensitivity, and intervention design constrain the interpretation of behavioural changes caused by erasure changes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.