acceptodds
Under review as a conference paper at ICLR 2027

What Do Linear Deception Probes Actually Detect?

Abstract

Linear probes trained on language-model activations can detect strategic deception across diverse settings, yet their performance varies substantially across tasks and layers, raising a fundamental question:_what do linear deception probes actually detect?_ We show that a widely used representation-engineering construction primarily captures an instruction-induced persona state. Its training pairs contrast honest and deceptive instructions while holding the assistant's continuation fixed and truthful. Controlled interventions reveal that probe scores are substantially more sensitive to instruction framing than to output truthfulness. Moreover, benign role-playing instructions account for much of the instruction-induced shift, indicating that the learned signal is not specific to deception. We introduce a matched-difference construction that contrasts true and false continuations while holding instruction framing fixed. Across 12 models from four families, the resulting probe shows reduced framing sensitivity and stronger content discrimination under previously unseen instructions and improves mean AUROC on four of five Liars' Bench tasks. Together, our findings identify persona-related confounding in standard deception probes and provide a counterfactual training approach that more directly targets belief-contradicting content.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.