acceptodds
Under review as a conference paper at ICLR 2027

Characterizing Rhetorical Misalignment in Decision-Making with Language Models

Abstract

Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether their assistance can amplify these biases and lead to harmful consequences. In this paper, we study rhetorical misalignment: a failure mode where a language model presents information in rhetorically inappropriate forms, thereby inducing suboptimal human decisions in a given decision context. We conduct a human-subject experiment using decision problems curated from the United States Medical Licensing Examination (USMLE). We observe an average harmful flip rate of 2.81% across models, where participants revise initially correct answers to incorrect ones after viewing responses generated by language models. Participants' rationales suggest possible mechanisms related to the language used by the models, consistent with cognitive biases including anchoring, authority bias, and loss aversion. Motivated by the observations, we develop a decision-theoretic framework that characterizes rhetorical misalignment through the expected utility gap between rational and behavioral use of the information provided by language models. To empirically measure rhetorical misalignment, we instantiate the framework using decision-makers simulated by language models. Overall, our findings reveal a safety concern previously underexplored in high-stakes domains: language models can be factually aligned yet still induce harm through its rhetorical presentation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.