Does Chain-of-Thought Follow the Computation? A Graph-Based Study of Reasoning Faithfulness
Abstract
Recent progress in large language models (LLMs) has renewed interest in whether the chain-of-thought (CoT) that a model writes is the real reason for its answer. Existing faithfulness metrics treat a CoT as a line of text: they cut it short or insert a mistake and check whether the answer flips (Lanham et al., 2023), or they track how the model’s confidence decays along the chain (Ye et al., 2026). Such metrics compare a CoT only with itself, but the computation behind many problems is a directed acyclic graph (DAG). We introduce metrics that compare a CoT with the exact computation graph G⋆ that a solver program produces for every problem. Order fidelity asks a simple question: does the model state each value in the same order that the computation actually needed it? Path-monotonic fidelity (PMF) asks a related but simpler question at the level of a single step. Propagation Fidelity (PF) comes from a direct experiment we call Dependency-Targeted Mistake Injection (DTMI): we deliberately corrupt one value inside the computation and then check, value by value, whether only the values that were supposed to be affected actually changed. We run both kinds of metrics on the same records, for five task families that cover shallow computation, dynamic programming, and P-complete computation, and seven open-weight LLMs. On the four families where order fidelity needs no correction, graph shape (depth, width, and treewidth) raises the R2 of order fidelity by 0.26, against 0.01–0.06 for the baseline metrics. On models that were not used in fitting, order fidelity reaches a held-out R2 of 0.38, against at most 0.06. Under DTMI the final answer changes in 71–82% of cases, so an answer-flip test would call every model faithful, yet PF is only 0.21– 0.26, because the error leaks into values that do not depend on the corrupted one. These findings suggest that CoT faithfulness reflects the structure of the underlying computation and should be judged against it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.