When Pleasing the User Compromises Care: AI Sycophancy in Caregiver-Mediated Healthcare
Abstract
Large language model (LLM) alignment often assumes that the user interacting with a model is also the primary beneficiary of its behavior. Yet in many consequential settings, one person directs the system while another bears the consequences, creating a potential conflict between accommodating the user and protecting the beneficiary. We study this problem in caregiver-mediated healthcare, where parents, pet owners, and caregivers make decisions for children, animals, and older adults. We focus on third-party sycophancy: whether an AI advisor shifts toward a caregiver’s preferred but clinically inferior action at the expense of appropriate care for the patient. We introduce an evidence-grounded benchmark spanning veterinary, pediatric, and older-adult care. Holding clinical facts fixed, we vary caregivers’ stated preferences and communication styles and evaluate both initial recommendations and multi-turn follow-up. Across frontier and open-weight models, stated inferior preferences cause substantial additional switching from appropriate to inappropriate care beyond failures observed during neutral follow-up, while communication style alone has a smaller effect. These failures accumulate across conversational turns, including on cases models initially answer correctly, and often carry moderate or severe potential harm. These results highlight a broader alignment challenge: when the user and beneficiary differ, accommodating user preferences can come at the expense of the party who bears the consequences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.