Plausibility Confounds Linear Truth Probes in Large Language Models
Abstract
Recent work has shown that large language models have linear directions in their internal representations that can distinguish true from false factual statements, raising the possibility of verifying factual claims directly from hidden states without relying on the LLM’s generated output. Despite this promise, we find that existing truth-probing datasets, methods, and evaluation protocols overlook an important confound: plausibility, namely false statements that are factually or semantically close to the truth. Building on the datasets and probing framework of Marks and Tegmark (2024), we mine near-miss falsehoods from the model’s own conditional distribution and construct a plausibility direction using only plausible-false versus random-false statements. With no true statements used for probe training, this false-only direction is strongly aligned with the original truth direction (cosine similarity 0.78–0.84) and, when evaluated on the true-versus-false task, it reaches 88–100% of the original truth probe’s AUC. These results suggest that truth and plausibility are strongly aligned in representation space. This observation also exposes an important weakness in standard truth-probe evaluation. Existing work primarily reports ranking-based metrics such as AUC on test sets containing true statements and relatively easy random falsehoods. While such metrics are threshold-free, they do not directly characterize the reliability of a probe used to classify a single input statement. We therefore complement standard AUC with fixed-threshold single-statement classification and explicitly stress-test probes on plausible falsehoods. Replacing random falsehoods with near-miss negatives causes a modest degradation in AUC, and a much larger failure in deployment-style classification: the false-positive rate on plausible falsehoods is 3–8 that on standard random negatives. Importantly, near-miss falsehoods mined by one reference LLM remain challenging for other model families, indicating that the effect reflects transferable structure rather than model-specific adversarial examples. Geometric analyses further show that the truth and plausibility signals are coupled in ways that cannot be sufficiently disentangled by simple linear or nonlinear post-probe processing without sacrificing useful truth information. Therefore, we propose a strategy of incorporating plausible falsehoods directly into the negative training set, which aims at discouraging reliance on plausibility during truth probing, and it improves truth probing effects across multiple factuality datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.