Brand-as-Memory: How Vision-Language Models Use Source Identity in Credibility Judgments
Abstract
When vision-language models (VLMs) judge a news source from a screenshot, how much does the publisher's identity matter relative to the article's evidence? We study this question with CueTrust, a controlled comparison of source cues and content evidence across seven VLMs. Brand identity has a larger effect than the tested content manipulation in five models; author bylines and in-text authority cues do not dominate content in any model. In Qwen2.5-VL-7B, the brand effect is 1.79 times the content effect, and changing the source can reverse the judgment for matched article content. Brand rankings also persist when the articles and questions are translated into Spanish and Chinese, although the credibility gap is smaller in Chinese. Linear probes and position-resolved activation patching show that source-related credibility is decodable before it strongly affects the readout: readout-position transfer rises at layers 18–20, with a similar late transition in InternVL3-8B. Adding a task-selected sparse-autoencoder direction reduces the brand gap by 41% while retaining 75% of content sensitivity; the same direction also shifts judgments of held-out outlets. These results connect source-dependent credibility judgments to late readout representations that can be causally manipulated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.