Measure Once, Steer with Many: Context Contamination and Composite Control for Persona Drift in LLMs
Abstract
Monitors and controllers of persona drift read a model's activations against a persona direction: the difference between the model's activations with and without the persona in the system prompt, extracted once before the conversation and then held fixed (the frozen direction). Re-extracted later in the dialogue, with the conversation so far as a prefix, this direction appears to rotate and shrink. We show that this movement is context contamination, not a moving persona. It follows how much new content the conversation adds: replacing the assistant's replies with another persona's barely changes it, replacing them with one repeated sentence slows it several-fold, and conversation length alone predicts it. The re-extracted direction is also the worse instrument: it does not track the judged persona within a conversation, and steering along it does worse. Steering harder along the frozen direction, however, degrades coherence and induces repetition. We therefore extract coherence and repetition as directions with the same recipe and steer along a weighted composite of all three, recovering most of the coherence and variety that pushing the persona costs. Against five published steering methods on four models, each at its best strength, some configuration of ours is at least as good on two of three measured qualities in nearly every setting. We also find that the more text accumulates, the harder the frozen direction is to express in other instructions' re-extracted directions, and that a malicious persona fades three to five times faster than a compassionate one. Our judges, blind to the method behind each reply, agree with people as closely as people agree with one another.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.