acceptodds
Under review as a conference paper at ICLR 2027

From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness

Abstract

Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input–output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as an internal concept grounding: Does a large language model's (LLM) CoT reasoning engage the same internal concepts that support the LLM's direct prediction, and do the shared concepts causally drive its answer? Concretely, we encode the prediction and CoT passes with a single shared sparse autoencoder (SAE), treating its latent features as concepts and thereby mapping both passes into a shared concept space. Within this space, we design three correlational metrics of concept-level alignment and a causal metric, , which measures the drop in answer probability when the shared concepts are ablated. Across five LLMs and four datasets, concept alignment is generally high, as indicated by the correlational metrics; yet these only identify which concepts are shared, not how much they causally contribute. fills this gap: causal faithfulness varies substantially with model depth, peaking at mid-to-late layers, and model scale reshapes the layer-wise profile. Moreover, causally important shared concepts are not always verbalized in the CoT. These dissociations suggest that faithfulness cannot be reliably assessed from surface-level or representational correspondence alone; assessing it requires causal tests of whether the internal concepts underlying a CoT actually drive the model's prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.