The Authority Trap: Isolating Persona-Induced Sycophancy in Vision-Language Models Across Safety-Critical Domains
Abstract
Professional users utilize Artificial Intelligence (AI) systems in high-stakes work environments using authoritative language. We investigate whether Vision-Language Models (VLMs) override their own visual judgments in response to this expert-toned language, a pattern we call Persona-Driven Sycophancy. In this study, we introduce a no-persona control condition (the identical false claim without any expert persona) to test if authority itself, or just confidence, causes this. We considered six widely-used VLMs across 900 responses spanning 50 images across 13 professional domains, from general scenes to specialized technical fields such as aerospace, civil engineering, and mining, each paired with a baseline, expert-persona, and matched no-persona control prompt. Open-source models exhibit substantial (37% to 99%) sycophancy, fabricating information to appease confident claims, while frontier models are comparatively robust (0% to 22%). Critically, isolating the persona variable, we do not detect a clear effect of expert framing (−2 to +4 percentage points) beyond the claim’s confidence at this sample size, a null result consistent with VLMs responding to confident assertion regardless of who states it. Sycophancy also changes a lot by domain, from 27.8% in botany to 69.4% in engineering. We also observe a preliminary over-refusal mode, where safety filters occasionally block legitimate queries in frontier models. We check our scoring method (an LLM judge) against human ratings and find 96.5% agreement (Cohen’s κ = 0.934). Overall, our results suggest VLMs may respond more to how confidently a claim is stated than to who is making it, a consideration for their deployment in safety-critical professional environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.