acceptodds
Under review as a conference paper at ICLR 2027

DPEB: Evaluating Persona Evolution in Multi-Session Dyadic Interaction

Abstract

Long-horizon conversational agents are increasingly expected to sustain a coherent persona while adapting to the relationship they build with a single returning user. The decisive question is temporal and dyadic: does the persona change when, and only when, the user's situation warrants it? Research to date has treated persona as a static target to be steered or a property to be monitored, leaving this question unmeasured. We introduce DPEB, a controlled benchmark for dyadic persona evolution: 1,200 fully reproducible scenarios, balanced by event type at the first evolution point, each pairing an agent persona with one user's relationship arc and labeling objectively the sessions in which a Big Five dimension should move and those in which it should not. Personas are carried by activation steering, measured by projecting the residual stream onto contrastive persona vectors, and scored on three axes (evolution responsiveness, stability, direction accuracy) against an oracle that updates at exactly the labeled sessions. Across the models evaluated DPEB yields a reframing result: static regimes score zero by construction and history-conditioned prompt-update baselines barely move, whereas a re-implemented Persona-Flow-style baseline moves readily but drifts, and holding the update rule fixed while varying only the evidence detector changes the outcome far more; measured on the evaluated scenarios, the two LLM-based gates reach 97-99% event recall yet false-trigger on 36-40% of ordinary (non-evolution) turns, while a prototype detector reaches 96% event recall at a 1.7% false-trigger rate. A lightweight detector trained on sentence embeddings combines both advantages (0.90 event-type accuracy at a 2.3% false-trigger rate) and, with the update rule held fixed, improves all three axes significantly over both deployed gates on two backbones, reaching 93-98% of the oracle's evolution responsiveness. We formalize this as a hierarchical view separating a slow, evidence-licensed trait trajectory from fast within-session adaptation, and we release the scenarios, the frozen user side, the measurement code and the detector diagnostics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.