Controls Before Claims: Auditing Truth and Context-Faithfulness Steering Directions
Abstract
Large language models can produce fluent answers that rely on false information or ignore the evidence in the prompt. We study truthfulness and context faithfulness through directions in activation space. We extract a truthfulness direction, a context-presence direction, and a faithful-answer direction, and we evaluate them on five models (Llama 3.1 8B Instruct and base, Qwen 3.5 27B, Qwen 2.5 7B, and Mistral 7B) against controls that the steering literature usually omits: norm-matched random directions, dose matching, layer sweeps, extraction-source ablations, and judge-scored informativeness. On every model, a context direction steers context-following under knowledge conflict far above random directions of the same size (up to +72 points on Llama 3.1 8B Instruct), with informative answers, and the effect is invariant to the extraction data. Which contrast works is model-specific: the faithful-answer direction dominates on Llama and Qwen 2.5, the presence direction on Mistral. The truth direction reads whether the supplied context is false before generation on every model, but steers truthfulness only when it is polarity-balanced and applied at the layers where a negation control shows it encodes truth (TruthfulQA MC1 0.51 to 0.59, above 30 random directions); at the layer the standard calibration selects it only suppresses answers. After whitening, the truth and context directions share no geometry. These results show that direction-based claims depend on controls and layers that are rarely reported; they do not show that any single direction is a reliability intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.