What Do Linear Deception Probes Actually Detect?
Abstract
Linear probes trained on language-model activations can detect strategic deception across diverse settings, yet their performance varies substantially across tasks and layers, raising a fundamental question:_what do linear deception probes actually detect?_ We show that a widely used representation-engineering construction primarily captures an instruction-induced persona state. Its training pairs contrast honest and deceptive instructions while holding the assistant's continuation fixed and truthful. Controlled interventions reveal that probe scores are substantially more sensitive to instruction framing than to output truthfulness. Moreover, benign role-playing instructions account for much of the instruction-induced shift, indicating that the learned signal is not specific to deception. We introduce a matched-difference construction that contrasts true and false continuations while holding instruction framing fixed. Across 12 models from four families, the resulting probe shows reduced framing sensitivity and stronger content discrimination under previously unseen instructions and improves mean AUROC on four of five Liars' Bench tasks. Together, our findings identify persona-related confounding in standard deception probes and provide a counterfactual training approach that more directly targets belief-contradicting content.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.