acceptodds
Under review as a conference paper at ICLR 2027

Same Task, Different Answer: Persona Conditioning Modulates Shared Task Mechanisms in Language Models

Abstract

Persona conditioning can change a language model's behavior on a task even when the task itself is held fixed, but it remains unclear whether these changes arise from persona-specific mechanisms or from modulation of ordinary task computation. We investigate this question across three instruction-tuned language models, five tasks, and four persona axes. Applying complementary damage and repair path patching to persona-conditioned correct–wrong transitions, we identify attention-head circuits and decompose each into a component shared across axes () and an axis-dependent remainder (). Circuits discovered in opposite causal directions converge preferentially on , and faithfulness, knockout, direct logit attribution, probing, and no-persona validation consistently implicate in task-related computation. Moreover, the bidirectional intersections retain strong task-functional signatures and much of the intervention effect while using substantially fewer heads. As an independent validation, in Qwen-7B SST-2 every persona-derived shared head falls within a separately discovered no-persona task circuit, whereas the remainder heads show much lower containment. The same structure is also behaviorally actionable: a signed intervention selectively improves persona-sensitive accuracy while preserving more than 99% of originally correct outputs. However, how strongly persona information couples to this shared computation depends on the task: the behaviorally stable ARC tasks show greater separation between task and persona information, whereas ETHICS, Safety, and SST-2 show stronger coupling within . Overall, these results show that persona-conditioned behavior can emerge from modulation of shared task-related mechanisms, not only from isolated persona-specific circuitry.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.