Knowing Whom to Trust: Reliability States and Information Use in Large Language Models
Abstract
Agents based on large language models (LLMs) increasingly make decisions using information supplied by people and other agents. Their choice of whom to trust shapes both the information they accept and the resources they entrust. Alongside studying whether agents themselves are trustworthy, an equally important question is how they learn whom to trust and which internal processes connect that experience to later decisions. We investigate this question across eight general-purpose LLMs spanning multiple model families and parameter scales. Models receive reports and verifiable feedback during interactions with four unfamiliar simulated participants whose reporting reliability is experimentally controlled, then allocate resources in a trust game and choose between conflicting recommendations. Participants with higher reporting reliability receive larger transfers and are selected more often when recommendations conflict. In Qwen3.5-27B, verbalizer readouts and linear probes distinguish participants with high versus low reporting reliability at the output and activation levels. Activation patching and steering shift recommendation-choice probabilities and expected transfers to the corresponding participants; the same steering direction affects both decision tasks. These findings link interaction experience to subsequent decisions through internal evaluations of reporting reliability, providing behavioral and causal evidence for how LLMs evaluate others and decide whom to trust.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.