A Theoretical Perspective on Why LLMs Get Lost in Multi-Turn Conversations
Abstract
Recent empirical studies show that LLM task performance can degrade when the same user information is distributed across multiple turns rather than presented at once (Laban et al., 2026). We study this phenomenon theoretically by modeling multi-turn conversation as closed-loop inference over a fixed user intent: the model repeatedly updates its interpretation, generates a response, and may reuse that response as evidence later. This feedback loop creates a fundamental asymmetry between the two sources of conversational evidence: user messages provide exogenous information about the true intent, whereas model responses are generated from the model’s current interpretation and therefore tend to reinforce it. We show that user evidence alone matches the corresponding single-turn error probability and becomes exponentially reliable as additional information arrives. Although model-generated evidence is Bayes-correct in a reference setting where user and model messages are generated from the same latent intent, in the closed loop it instead reinforces the current estimate, whether correct or incorrect. This competition creates distinct recovery and lock-in regimes: sufficiently strong feedback can make errors persist indefinitely, producing a non-vanishing multi-turn performance gap and coexistence of persistently correct and incorrect conversations. We further characterize how pruning or delaying model-history reuse and increasing the relative influence of user evidence mitigate these effects.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.