acceptodds
Under review as a conference paper at ICLR 2027

The Truthfulness Spectrum: How General Are Truth Directions?

Abstract

Large language models (LLMs) have been reported to linearly encode truthfulness, but recent work questions whether there exists a domain-general truth encoding. We reconcile these findings through the truthfulness spectrum: truth-related linear directions lie on a spectrum of generality, from broadly domain-general to narrowly domain-specific. To characterize this spectrum, we evaluate probe generalization across five truth categories (definitional, empirical, logical, fictional, and ethical), sycophantic and expectation-inverted lying, and existing honesty benchmarks. We find that linear probes generalize across most domains, while sycophancy transfer is near-chance and expectation-inverted transfer is often reversed. Yet joint training on all domains yields a single linear direction with strong held-out performance; poor transfer between domains thus does not preclude a domain-general detector. Prior evidence against a shared truth direction also includes comparing probe weights, but this too can obscure shared signals: nearly orthogonal probes can behave similarly on the same data, and cosine similarity of probe weights explains only about half the variance in cross-domain transfer (R^2=0.56). To further characterize the spectrum constructively, using novel concept-erasure methods, we isolate directions that are highly domain-general, domain-specific, or shared across certain domain subsets. These directions serve different causal roles. In steering experiments, several domain-specific directions increase the correct answer margin, while the domain-general direction is null or harmful despite strong detection performance. Finally, we show the spectrum is reshaped by post-training, which reduces alignment between sycophancy and other truth types. Together, our results reconcile prior seemingly conflicting findings and further characterize truth directions on a spectrum, with directions of varying generality coexisting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.