acceptodds
Under review as a conference paper at ICLR 2027

Knowing Is Not Saying: How Language Models Decouple Truth from Deceptive Reporting

Abstract

Large Language Models can produce incorrect reports under deceptive instructions even when the truthful answer remains internally available, raising a fundamental question: how can knowledge and report diverge within the same model? We study instructed lying as a controlled setting for this problem and refer to the resulting separation between available knowledge and reported behavior as knowledge–report decoupling. Layerwise decoding shows that deceptive instructions reshape the answer space through exclusion rather than substitution: they broaden incorrect competition while the truthful target retains predictive support and remains linearly decodable deep into answer formation. Per-layer transcoder (PLT)-based circuit tracing further shows that truthful and reported incorrect targets share highly similar input-side attribution but diverge substantially in downstream target-directed influence, revealing a distributed target-dependent pattern rather than a dedicated internal state that partially transfers to broader deception-related settings. Feature interventions further distinguish answer-specific from shared report control: full downstream effects are most effective for individual questions, whereas mediated influence is strongest across questions, suggesting that shared report control lies in the transformation from knowledge to report, rather than at the output endpoint.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.