acceptodds
Under review as a conference paper at ICLR 2027

Decorative Rationalization in Large Language Models

Abstract

Many modern uses of LLMs involve answering questions over information supplied at test time – diagnosing a patient from test results, reviewing a contract for risky clauses, or debugging code from an error trace. In these tasks, the model returns an answer along with rationale that people could judge to decide whether to rely on the answer. In this paper, we reveal that models produce decorative rationale, citing parts of the context as reasons for an answer that does not depend on them. We formalize this failure through the probability that removing a cited part leaves the answer unchanged, and estimate it through causal intervention on instruction-following and reasoning models. Our results reveal that much of the cited information is used decoratively. On average, removing a cited part leaves the answer unchanged 45 to 87% of the time. Decoration survives every condition we test. It persists when the model reasons before answering and when the model cannot answer without the context. It also persist when we strip the decorative context, because the model re-decorates with whatever remains. LLMs thus rationalize, creating a harmful illusion that may only be fixable at training time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.