CRISP: Validity Criteria for Intervention-Based Chain-of-Thought Faithfulness
Abstract
Chain-of-thought faithfulness is measured by intervening on a model's reasoning and recording the answer change. That contrast is computed under teacher forcing, so it is interpretable only where the replay reproduces the model and the intervention changes what the chain implies. Neither condition is usually tested. We introduce CRISP, a Controlled Reasoning Intervention with a Style-matched Perturbation. One generator writes a semantics-preserving control and a targeted semantic perturbation at matched length and style. The accuracy drop divided by the control's accuracy above chance has reference points at 0 and 1 and needs no baseline. Replay validity requires teacher-forced replay to reproduce the model's own answer, and intervention validity requires the perturbation to lower an evaluator's answer recovery. Both apply to any teacher-forced intervention. Across 15 models and 9 benchmarks, 30 of 135 pairs fail replay validity and 28 more fail intervention validity. Four checkpoints fail replay validity on clinical questions and pass on mathematics. The truncation baseline is negative on 6 of their pairs. Where intervention validity fails, the perturbation changed the recoverable answer on 30% of the checked items per pair against 55% on verified pairs, and where it did, the effect was 0.605 against 0.749. Held-out CRISP is 0.337 higher after CRISP-selected training traces. The criteria add 1 teacher-forced condition and 3 evaluator calls per checked item, so a study can establish that its numbers are interpretable before using them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.