SIFT: Robust Meta-Faithfulness Verification of Chain-of-Thought Reasoning Under Distribution Shift
Abstract
Chain-of-Thought (CoT) faithfulness detectors are widely used to audit reasoning models, yet a detector is itself a predictor whose verdicts are treated as stable properties of the model. We ask whether a faithfulness detector is faithful to itself under distribution shift. We formalize **meta-faithfulness** as an invariance principle: a valid detector must return identical verdicts on traces that differ only by transformations preserving ground-truth faithfulness. We prove three results: a behavioral indistinguishability theorem showing that no detector operating on intervention-response profiles can separate faithful from epiphenomenal mechanisms with identical signatures; an invariance-violation lower bound showing that any detector relying on shift-sensitive features must violate invariance at a rate independent of its in-distribution accuracy; and an asymptotic certified selective-risk guarantee enabling abstention with high confidence. We operationalize this principle in FaithShift, a stress-test protocol spanning ten shift axes, and propose **SIFT** (**S**hift-**I**nvariant **F**aithfulness **T**rajectory Detector), a hidden-state trajectory detector trained with cross-environment invariance objectives and certified abstention. Across 14,996 traces, four reasoning domains, and eight model architectures, three findings emerge. *First*, transfer collapse is real: all existing detectors exhibit gaps AUROC. *Second*, the dominant bottleneck is sampling stochasticity, not shift sensitivity: over 80% of detector instability stems from random seed variation rather than distribution shift, falsifying our preregistered hypothesis that shift-attributable invariance violations would exceed 0.25. *Third*, SIFT reduces shift-attributable instability, but its advantage vanishes against ensembles: SIFT cuts invariance violations by 64% over the best single-seed baseline, but a four-seed ensemble of any detector narrows the margin to 0.01, and at matched coverage the two are statistically indistinguishable (). SIFT also requires a 51% abstention rate that exposes a robustness-usability trade-off. Cross-model transfer degrades along a clear hierarchy, within-family cross-family open-weight open-weight to API, which multi-model training on 3-4 models partially closes. We therefore position this work as a framework for auditing auditors and a diagnosis of the actual barrier to reliable faithfulness detection: not distribution shift, but detector variance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.