One Draw Is Not Enough: Resampling-Based Labels for Chain-of-Thought Unfaithfulness
Abstract
Training and evaluating chain-of-thought monitors require reliable labels of whether reasoning reflects what influenced a model’s answer. Cue-based evaluations label a trace unfaithful when the model switches to a suggested answer without crediting the cue, but a single switch can also reflect sampling variability. An existing correction estimates sampling noise in aggregate but cannot validate individual labels. We introduce CueBall, an open-source pipeline that supplements these labels with repeated samples. It measures whether cue-following persists and how each reasoning trace credits the cue. We evaluate five open-weight reasoning models on four multiple-choice datasets with eight cue styles, at temperature 0.7 on questions screened for consistent uncued answers. Among switches from a correct baseline answer to an incorrect cued answer, about 40% of traces labeled unfaithful fail our persistence criterion: the cued answer must recur in at least three of four fresh samples. Across switches to the cued answer, the aggregate noise estimate is closer to the fraction with no observed recurrence than to the fraction failing this threshold. Failure to meet the threshold does not establish sampling noise, and persistence does not establish individual cue causation. We therefore recommend pairing trace-level faithfulness labels with evidence of behavioral persistence. To support this, we release 20,000 traces judged unfaithful or incoherent from 9,997 persistent model–question–cue pairs covering 2,610 questions, with labels, prompts and validation artifacts for monitor and probe research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.