acceptodds
Under review as a conference paper at ICLR 2027

When Do Deception Probes Transfer? Wrongness Controls and Source Specificity

Abstract

A deception probe may transfer better with more source labels because it learns deception-specific information, or because it becomes better at detecting ordinary wrong answers. We distinguish these explanations by inducing five wrong-answer policies while holding the task, model and item difficulty fixed. Each transfer curve is compared with a wrongness control that scores the same direction on honest errors. On ARC with Llama-3.1-8B, difference-of-means transfer from a hidden-objective source rises faster than the control. The supervised contrasts have the same sign but are unresolved after full refitting. A question-disjoint Gemma-2-9B replication resolves the scheming contrast for both estimators. On MMLU, by contrast, transfer and wrongness rise together. The diagnosis changes across task–model configurations, including configurations using the same benchmark. Even at larger label budgets, targets are usually detected best by source-matched directions. These results motivate reporting a powered wrongness control when interpreting probe-transfer curves.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.