SemFaith: Semantic Geometry for Chain-of-Thought Faithfulness Evaluation
Abstract
Can a reasoning trace reveal when a model conceals its use of a misleading hint? We investigate this question through the geometry of sentence embeddings, using BonaFide's hint-acknowledgment labels as an observable test case. Our hypothesis is that how a trace's steps spread across semantic regions, how abruptly they shift topic, and how far they stray from reference centroids carry information about whether the hint is acknowledged. SemFaith segments a chain of thought (CoT) into steps, embeds them with a fixed encoder, and summarizes these measurements at three clustering resolutions; a regularized linear head maps the resulting 24 features to a score without accessing the audited model's hidden states or querying an auxiliary judge. Across five held-out Instruct models (662 CoTs), SemFaith attains 0.885 mean fold AUROC. On 67 within-model pairs closely matched for step and word counts, its AUROC is 0.788, compared with 0.533 for a length-only control, showing signal beyond gross trace length. A supervised TF-IDF baseline reaches 0.991 on the full held-out evaluation, exposing strong lexical cues in this benchmark. These results identify a compact geometric predictor of hint concealment whose signal is not explained by gross trace length, while delimiting what BonaFide can establish about general reasoning faithfulness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.