CHAIN-OF-THOUGHT RELOCATES THE FACTUALITY SIGNAL: A LOCKED-PROTOCOL AUDIT AND CLAIM-PAIRED RECOVERY OF INTERNAL PROBES
Abstract
Internal-state factuality probes do not lose their signal under chain-of-thought:the signal moves. Across four model families, under a single locked protocol that varies labeling budget, training style, and feature pipeline orthogonally, we make three measurements. First, the direct-to-CoT transfer loss is small relative to protocol choices: the feature pipeline alone shifts CoT-side scores by an order of magnitude more than the style change itself, and target-style labels close most of the residual. Second, the residual is structured internal geometry: a claim-paired affine map recovers it, shuffle-paired controls collapse to chance, generic unsupervised alignment fails on every family, and label-free logit distillation from the same pairs nearly matches the full map – recovery is a label-transfer phenomenon that runs through the claim-paired internal geometry rather than through surface form. Third,natural false claims are measurable without synthetic corruption, and the labels themselves are auditable: adjudicating every generation shows that automatic labels conflate asserted errors with unasserted truncations, and a strict audit retracts two of our own label regimes as artifacts – yet the detection signal survives label cleanup.Protocol and labels determine what a factuality probe measures; the audit itself is the transferable finding, with claim localization and corpus-specific generalization the open frontiers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.