Probe-Driven Adversarial Co-Evolution for Distinguishing Hallucination and Deception in Large Language Models
Abstract
While large language models (LLMs) are increasingly used for question answering and decision-making, they can still produce incorrect responses. These errors may arise from (1) hallucination, where the model unintentionally produces incorrect content, or (2) deception, where the model intentionally produces false or misleading information. Existing methods mainly rely on LLM outputs or follow-up interactions to identify errors, but usually do not determine whether an error comes from hallucination or deception. Distinguishing the two is essential for diagnosing model failures and applying targeted mitigation strategies. We argue that effective deception detection should actively uncover diverse deceptive patterns, enabling the detection to generalize across varying deception strategies. In this paper, we propose a black-box method that combines probe interaction with adversarial strategy co-evolution to distinguish hallucination from deception in LLMs. The probe interaction module collects behavioral evidence by observing how the model explains, verifies, and revises its answer under interrogation. The adversarial optimization module optimizes a lie generator and a discriminator through a multi-round competition, allowing the detection strategy to cover broader deceptive patterns and target deception-specific behavioral cues. Experiments across multiple LLMs and datasets show that our method outperforms strong baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.