: Revealing Truth-Direction Sensitivity Loss under Sycophantic Pressure
Abstract
Language models can retain decodable truth-related information while their answers change under user pressure. Assessing the robustness of these representations requires measuring how strongly answers respond to the same information. We introduce Probing–Sensitivity Inspection (PSI), a protocol that pairs a frozen linear probe with a two-sided intervention along the same fitted truth direction. On single-token true/false tasks with Qwen2.5 (7B, 14B, 32B) and Gemma-2-9B, PSI reveals three findings. First, decodability–sensitivity dissociation: probe accuracy remains at 1.000 across the tested prompt conditions and cue levels, while intervention responses vary and reverse under instructed inversion; Qwen2.5-7B then answers truthfully on 0.133 of items. Second, sensitivity attenuation under preserved accuracy: under honest system prompts, sensitivity declines with sycophantic pressure in both model families while truthful-answer accuracy remains high. At the first cue level on Qwen2.5-7B, the intervention effect falls by more than one fifth while accuracy is 0.933. Sycophantic cues produce larger reductions than length-matched neutral prompts. Third, projection-based truth recovery: under preference-conditioned prompts, amplifying each item's existing projection raises truthful-answer accuracy from 0.700 to 0.883 on Qwen2.5-7B and from 0.500 to 0.850 on Gemma-2-9B, with little change under norm-matched random directions. PSI reveals pressure-dependent changes in truth-direction sensitivity that remain hidden behind stable probe accuracy and strong answering performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.