acceptodds
Under review as a conference paper at ICLR 2027

What does lie-detection accuracy measure? An Identification Audit of behavioural LLM deception benchmarks

Abstract

**Accuracy on instructed-lie benchmarks does not identify deception: the intervention that creates the lie also changes instruction-following behavior**, so its accuracy cannot be read as evidence of deception **as opposed to instruction-induced behavior**. **We do not ask whether LLMs can detect deception; we ask whether current benchmarks provide evidence that they do.** The reason is structural: **no experiment that intervenes only on the instruction can separate a detector response mediated by deception from one mediated by instruction-following**, both being descendants of that instruction. A deception-attribution reading needs , which (the *instruction*'s effect, and not a magnitude accuracy estimate) neither bounds nor signs, so no gain in accuracy or scale resolves it. **This is an identification problem, not a generalisation one: generalisation can be perfect while identification fails.** **Our contribution is an operational identification audit for behavioural LLM benchmarks, together with empirical evidence that the common instructed-lie paradigm produces large, reproducible attribution failures across three detector paradigms, and that benchmark design can determine which behavioural signal appears to constitute "deception detection".** The audit has five parts: **one necessary construct-validity test no instructed design can supply**, three alternative-explanation diagnostics and one robustness test. Fixed elicitation itself is not new (trained belief-verified organisms have it), but not as the *construct-validity criterion* for reading instructed-benchmark accuracy. Across three detector paradigms, equalising the prompt **leaves no detectable above-chance discrimination on any of the six targets** (97.0% **43.3%**), and **the accuracy the benchmark is scored on is reachable with no learned features at all**: a parameter-free lexical rule reaches 69–80%, *surface-accessible accuracy*, not a decomposition of the detector's signal. Because **no public release we audited supplies all five**, we built materials that do: with deception operationalised independently of the detector's input, prior work's own battery separates operationally verified deception on three of five targets (two underpowered nulls), of which **a blinded re-run replicates two and a pre-registered covariate audit withdraws one, so the standing count is one of five**, with **one of five *new* targets positive**. That demonstrates that criterion-4-valid materials can produce a deception-associated signal, **not an estimate of the prevalence or reliability of deception detection**. **Read instructed-benchmark accuracy as evidence about the intervention, not about deception detection**: the canonical benchmark fails to identify deception, while a fixed-elicitation design reveals *target-dependent* deception-associated signals (still not ), **rung 4 of five**. **Scope: English; open-weight models 3B–70B, plus seven frontier-scale targets from seven organisations; instructed-roleplay evaluations except where criterion 4 requires otherwise.**

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.