Interpreting and Steering Trajectory Entrainment Behavior in Goal-Oriented Conversations
Abstract
An emerging characteristic of conversational Large Language Models (LLMs) involving longitudinal dyadic interaction is the asymmetric drift towards attractor states. Interpreting this drift is crucial for both sustaining intended behavior and preventing unintended behavior in functional long-horizon deployments. Building on the idea of using dyadic self-play trajectories as attractors, we formalize this gradual adaptation behavior as and extend it to include turn-to-turn persistence, or . Across four LLMs and four conversational settings, we observe entrainment over 10 turn pairs along five dimensions: style, semantics, emotion, task, and personality. We find personality and task exhibit stronger linear separability than style and semantics in deeper layers, and this separability persists into later turns. Turn-to-turn persistence appears to vary across settings, with emotion exhibiting stronger inertia in personal advice conversations than in less affective settings. We further test whether norm-normalized contrastive activation addition can steer these dynamics. Steering along attunement directions induces persistent convergence and divergence in conversational trajectories while preserving fluency and grammaticality. Together, our findings characterize conversational entrainment and its persistence across models and settings, and show how activation steering can modulate these dynamics. These results demonstrate how steering can be used to correct temporally drifting LLM output along various goal-oriented attunement axes in autonomous conversational agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.