LLMs Judge You Harshly When You Speak Figuratively
Abstract
LLMs are increasingly used to judge people and their work, from moral advice to scientific paper review. We show that these judgments become harsher when the question is phrased figuratively. Given a moral scenario, we rewrite only the question as metaphor, hyperbole, simile and sarcasm, verified by human annotators to preserve the same meaning. We then evaluate nine models from six families, including both open-source and proprietary state-of-the-art models. All nine models offer a harsher judgment score, with GPT-5.6 Sol judging the user more harshly 17.5% of the time, comparable to human performance. However, humans exhibit a significantly lower flip rate (i.e., changing their judgment polarity from positive to negative or vice versa), while LLMs, especially open-source models, flip their decision by up to 11.1% of the time. This effect extends to mathematical reasoning, truthfulness QA, scientific paper review and resume review. We further conduct analysis to show that figurative questions draw the model's attention away from the scenario, providing cues for the judgment shifts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.