acceptodds
Under review as a conference paper at ICLR 2027

Evaluating stability and generalization of truth and safety probes in LLMs

Abstract

Linear probing is widely used to investigate how language models represent concepts such as truthfulness and safety, and has been proposed for monitoring and steering model behavior. Prior studies report contradictory results: some find shared linear directions that transfer across datasets, while others report poor cross-dataset robustness or stability. However, these studies differ in language models, datasets, prompt formats, and evaluation procedures, making their conclusions difficult to reconcile. Guided by the Predictability–Computability–Stability (PCS) framework for veridical data science, we treat these experimental choices as reasonable perturbations within a unified study and ask which conclusions about linear probing remain stable or change across them. Our study spans twelve datasets across medical, factual, statistical, safety, and social-reasoning domains, as well as different language model layers, pretrained architectures, prompt formats, input feature types, probe classification algorithms, and training-domain coverage. We find that truth- and safety-related activation representations across domains are better described by hierarchical linear probes that jointly model shared structure alongside dataset-specific variation over flat linear or nonlinear probes. In assessing the role of distribution shift, we find that probes trained on data including target distribution samples consistently outperform those trained out-of-domain, while aligning prompts to a standard True/False format improves cross-domain transfer. Nonlinear probe architectures and alternative feature representations either match or underperform linear probes trained on raw activations. Finally, predictive performance and steering effectiveness are distinct properties: the directions that most strongly steer the evaluated models align with maximal activation variance rather than discriminative power.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.