acceptodds
Under review as a conference paper at ICLR 2027

How Language Models Respect Role-Specific Cognitive Boundaries: A Causal Account

Abstract

Language models know far more than the roles they portray. Faithful role-playing requires them to respect a role's cognitive boundary—the range of answerability supported by its profile. Although existing work has evaluated and improved adherence to these boundaries, it offers limited insight into how role information constrains answering internally. To investigate this question, we construct a controlled role–query dataset in which the same query lies within one role's cognitive boundary but beyond another's. Through causal interventions on these counterfactuals, we uncover a representation-level process in which the model encodes the role profile and uses it to condition the processing of the current query. When the query lies beyond the role's cognitive boundary, the model reads the conditioned query representation and progressively writes an abstract conflict signal where the response is predicted. This signal remains effective when transferred across role–query contexts and is translated into a specific response only in later layers. Our findings provide causal evidence that models internally assess queries against role-specific cognitive boundaries and represent conflict as a shared signal, clarifying how role profiles constrain what they present as known.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.