Diagnosing and Mitigating Sycophancy in LLMs via Attribution-Guided Activation Steering
Abstract
Sycophancy refers to the tendency for LLMs to match user beliefs at the cost of factual correctness, thereby undermining model reliability. Prior work on evaluating sycophancy in LLMs aims to assess whether a model's output matches an authority's claim, but cannot reveal which part of the prompt drives this sycophantic behavior. To bridge this gap, we investigate the relationship of sycophantic responses with an authority's credentials, their assertive claim, and the problem statement. We introduce the Attribution Share Index (ASI), an Integrated Gradients-based token attribution method, which measures the degree to which a model's decision is driven by authority-related text. Through experiments across five models and 30 test configurations, we find that sycophantic responses consistently direct more attention toward authority tokens than resistant ones. Moreover, our token attribution method reveals that for the sycophantic cases, the claim asserted by the authority receives more influence than the authority's credentials. Building on these findings, we use attribution-guided contrastive activation steering to mitigate LLM sycophancy by helping the model resist wrong claims while largely preserving accuracy when the authority is correct. We construct a steering vector by comparing the model’s internal states before and after masking three tokens that IG selected from the authority's claim. This vector can be applied at inference time without retraining. Together, our results show that token-level attribution can both explain what drives sycophancy and inform a practical intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.