Can We Trust Agents’ Chains of Thought? Measuring and Detecting Trace-Grounded Inconsistency
Abstract
LLM agents’ chains of thought are often used to understand why they take particular actions and produce particular responses. Existing studies of agent performance focus largely on task success, which does not reveal whether an agent's stated reasoning is consistent with its actual behaviors and observations. A key remaining question is therefore to what extent these reasoning traces provide a reliable account of an agent's actual behavior. We study trace-grounded inconsistency, where statements in an agent's chain of thought or response are inconsistent with its execution trace or its available evidence. We develop a taxonomy of such inconsistencies covering action commitments, execution, evidence-grounded claims, and reported outcomes. We find that these inconsistencies widely spread across all six agents evaluated, occurring in 27.2% of the collected conversations. Using this taxonomy, we construct a human annotated dataset of trajectories from six agents across three sandboxed application domains and analyze how different forms of inconsistency arise across agents and tasks. We further represent each trajectory as two structured sequences constructed from CoT and execution traces. Based on this representation, we train a model with reinforcement learning to generate executable checkers for automatically identifying inconsistencies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.