Training Honesty Adapters To Detect Deception In LLMs
Abstract
Activation probes trained on off-policy data are a cheap way to monitor language models for deception, but they often fail to generalize to on-policy deception outside their training distribution. We instead train an honesty adapter, a LoRA stacked on a model, to report whether the model exhibited a hidden behavior or deception in a transcript, using on-policy transcripts from 533 model organisms of Shenoy et al. (2026). Scoring a forced admission against a forced denial gives a continuous monitor score. Honesty adapters for Qwen3-14B and Llama-3.3-70B detect deception by held-out organisms with near-perfect AUROC and generalize to auditing organisms trained with a different recipe: on AuditBench, the Qwen3-14B adapter reaches 0.90 AUROC, against 0.74 for a probe trained on the same data and 0.66 for a probe trained on off-policy deception. On off-policy benchmarks of lies, reward hacking and hard-coded test cases, the adapters match deception probes in AUROC. Asked follow-up questions about a transcript, an honesty adapter also describes the organism's hidden behavior more accurately than an introspection adapter trained on the same organisms, and comparably or better without a transcript. Both on-policy training data and adapters in place of probes thus help deception monitors generalize.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.