acceptodds
Under review as a conference paper at ICLR 2027

The Condescension Trap in Language Models

Abstract

We introduce the condescension trap, a failure mode in which a language model treats some users as less capable than their request indicates, while appearing considerate. It parallels elderspeak, the simplified and reassuring speech that people direct at older adults and that recipients experience as patronizing. We show that preference optimization amplifies this behavior because the evaluator rewards it and the users who bear its cost are absent from the feedback loop. We test each link of this loop with age as the user attribute, on everyday requests whose answer does not depend on age. First, when the user states an age of 79 rather than 45, six language models change how they address the user: all six refer to the user's age more often, by 8 to 27 percentage points, with smaller increases in referrals to family and in reassurance. Second, LLM judges reward the change: with the answer pair held fixed and only the age told to the judge varied, three judges prefer the answer written for the 79-year-old by 18 to 28 points more when told the user is 79. Third, training on such a judge amplifies the change: three rounds of online DPO raise references to the user's age from 30% to 67% of answers to 79-year-olds, while hiding the age from the judge keeps them near 30%. Finally, adults 65 and over, comparing the same pairs blindly to condition, do not prefer the rewarded answer and rate it as slightly more likely to treat the reader as less capable. Withholding a task-irrelevant attribute from the judge prevents the amplification; grounding the judge in a content rubric reduces it without removing it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.