acceptodds
Under review as a conference paper at ICLR 2027

From Hidden States to Confessions: Auditing Deception in Reasoning Models

Abstract

As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern. Existing deception monitors typically either score visible transcripts or derive scalar probe scores from representation vectors, leaving little inspectable evidence about why a response is suspicious. We introduce StateWitness, an audit decoder for deception monitoring. It reads a target model's hidden states to answer audit queries and produce confession-style reports of potential failures or violations. We evaluate generalization beyond the decoder's audit training data across two target reasoning models and seven deception datasets. StateWitness achieves 0.916 mean AUROC, with relative gains of 11.6% and 25.0% over the best black-box and activation-probe baselines, respectively, under the same evaluation protocol. When combined with existing monitors, StateWitness reduces missed deceptive examples in simple threshold ensembles. Beyond scalar detection, the decoder returns query-level answers, schema reports, and token- or sentence-level evidence traces for human inspection. We envision StateWitness as one monitoring layer in a broader defense-in-depth strategy. Code is available at https://anonymous.4open.science/r/StateWitness-KX42/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.