acceptodds
Under review as a conference paper at ICLR 2027

DECOR: Auditing LLM Deception via Information Manipulation Theory

Abstract

Large language models (LLMs) can deceive users by subtly manipulating truthful information—omitting key facts, shifting focus, or obscuring meaning, making deception auditing essential. Existing black-box deception auditing methods typically make coarse-grained, response-level judgments without a theoretical basis for identifying the specific information involved in deception or how that information contributes to it. We propose DECOR, a multi-agent framework that operationalizes information manipulation theory (IMT) for deception auditing, which characterizes deceptive communication along four dimensions of information manipulation. DECOR decomposes input contexts into atomic information units, evaluates each unit against the response across these dimensions, and aggregates the resulting manipulation profiles into a global deception index. To enable reliable deception evaluation, we establish human reference labels for 745 LLM responses on DeceptionBench. Experimental results show that DECOR outperforms existing methods in both single-turn and multi-turn deception auditing, achieving 0.948 and 0.657 AUROC, respectively. Across 15 auditor backbones, incorporating elicited thoughts consistently improves DECOR, demonstrating their value as a complementary signal for deception auditing. Our analysis further shows that LLM deception more often involves omitting relevant information and diverting attention than fabricating or distorting facts. Our project is available at https://anonymous.4open.science/r/DECOR-F866/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.