When There Is No Single Right Answer: Do Sycophantic Judgment Changes Persist?
Abstract
Evaluating large language models for decision support requires assessing sycophancy in tasks without reference answers. We introduce SOLIPS, a binary-choice protocol for evaluating five models on eight benchmarks under explicit judgment criteria. It directly measures changes toward the user's requested alternative under repeated, unsupported pressure and their persistence at reevaluation. We retain the full conversation history during reevaluation and do not add new pressure statements. Under Bare rebuttal, the evaluated models show higher rates of judgment changes that persist at reevaluation on selected medical and legal tasks without reference answers than on corresponding tasks with reference answers. For Claude Opus 5, the corresponding rates are 72.7% on PrinciplismQA and 24.8% on MedQA-USMLE. This evaluation provides an empirical basis for assessing the reliability of LLM-based decision support under unsupported user pressure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.