acceptodds
Under review as a conference paper at ICLR 2027

Forecasting Persona Drift in Large Language Models Before Generation

Abstract

Language models used in continuing agent tasks may need to maintain an assigned persona over many turns. Detecting persona drift only after a response has been generated leaves no opportunity to inspect the risk before that response. Yet change alone is not evidence of drift: new task information can justify a different judgment without changing the assigned behavioral principle. We study whether sustained persona drift can be forecast before the next response is generated. Our method learns the internal state expected when the persona is maintained under the current dialogue and task conditions, together with its normal variation. It combines deviations from this stable behavior reference with recent state history and available dialogue to predict future drift. Experiments cover three sizes of Qwen2.5-Instruct and Gemma-3-4B-it. The method improves prediction over text and raw state baselines. Under a requirement to warn one full turn in advance, it reduces false warnings on justified updates while meeting the required recall and overall false warning limits. The reduction remains under a separate behavioral assessment and changes in dialogue order and wording, but reverses when relevant task information is unavailable. These results support monitoring persona stability before generation while allowing reasonable changes in judgment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.