Human History Helps-But for How Long? Tracking LLM Dialogue Continuation Across Turns
Abstract
Can human response history keep language-model continuations close to recorded human responses-and for how long after new human responses stop being supplied? We introduce a controlled dialogue-replay protocol that fixes one participant’s recorded utterances while a language model generates the target participant’s response at each eligible turn. Five conditions vary the target-speaker history available to the model: no recorded human responses, the first one, two, or five human responses followed by model-generated responses, or all previous human responses. Earlier human responses remain in context after each finite prefix ends, while the current human reference is never exposed. We evaluate five language models across informal conversation (MaiChat), emotional support (ESConv), and persuasion (PersuasionForGood), with over 630,000 model-generated continuations saved before matched-window filtering. MaiChat and ESConv use dialogue history alone, whereas PersuasionForGood additionally includes fixed task context without strategy labels. We measure turn-level semantic distance to recorded human continuations and estimate condition-dependent trajectories over matched dialogue windows using mixed-effects models. Across all 15 model-dataset analyses, full human history has a lower fitted window-averaged semantic distance than the condition without recorded human response history, although its fitted response-step trend remains positive. A five-response human prefix retains a detectable benefit over that condition after model-generated responses enter the history in ESConv and PersuasionForGood. Yet continued human responses provide an additional detectable benefit over that prefix at one or more post-prefix positions for all five PersuasionForGood models and four of five ESConv models. Effects in MaiChat are less precise and more model-dependent. These findings distinguish the carryover of early human responses from the benefit of continually supplying new ones. Our protocol measures fidelity to recorded continuations under fixed interlocutor utterances, not general response quality, task success, or fully interactive dialogue.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.