acceptodds
Under review as a conference paper at ICLR 2027

Controls Before Claims: Auditing Truth and Context-Faithfulness Steering Directions

Abstract

Large language models can produce fluent answers that rely on false information or ignore the evidence in the prompt. We study truthfulness and context faithfulness through directions in activation space. We extract a truthfulness direction, a context-presence direction, and a faithful-answer direction, and we evaluate them on five models (Llama 3.1 8B Instruct and base, Qwen 3.5 27B, Qwen 2.5 7B, and Mistral 7B) against controls that the steering literature usually omits: norm-matched random directions, dose matching, layer sweeps, extraction-source ablations, and judge-scored informativeness. On every model, a context direction steers context-following under knowledge conflict far above random directions of the same size (up to +72 points on Llama 3.1 8B Instruct), with informative answers, and the effect is invariant to the extraction data. Which contrast works is model-specific: the faithful-answer direction dominates on Llama and Qwen 2.5, the presence direction on Mistral. The truth direction reads whether the supplied context is false before generation on every model, but steers truthfulness only when it is polarity-balanced and applied at the layers where a negation control shows it encodes truth (TruthfulQA MC1 0.51 to 0.59, above 30 random directions); at the layer the standard calibration selects it only suppresses answers. After whitening, the truth and context directions share no geometry. These results show that direction-based claims depend on controls and layers that are rarely reported; they do not show that any single direction is a reliability intervention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.