Where Did We Go Wrong? Localizing Errors to Repair Multi-Turn Conversations
Abstract
Large language model (LLM) assistants are increasingly used in multi-turn conversations, where they must progressively infer a user's intent from partial information revealed over time. This makes them vulnerable to early misinterpretations that can rule out the true intent and persist across turns, ultimately causing the interaction to fail. Recovering effectively from such failures requires identifying which earlier interpretations need revision while preserving the information that remains valid. We introduce Loci (Localization Of Corrupted Inferences), a framework for probabilistically localizing such errors and selectively repairing the conversational state. Given an observed failure, Loci reasons over possible user intents and assigns each prior update a posterior probability of having excluded the true intent. These probabilities rank likely error locations and quantify confidence in the localization. We then use this confidence in a repair policy that decides whether to retain, verify, or remove suspected updates, enabling selective repair of the conversational state. To evaluate Loci's ability to identify corrupted inferences, we first test it in a controlled setup with injected errors. We find that Loci improves localization over the baselines, with its advantage becoming larger when multiple errors occur. We then evaluate it in free-form conversations, where errors arise naturally, and measure the additional interactions needed to complete the task after an initial failure. Using localization confidence to guide repair, Loci reduces mean recovery cost by 46–79% relative to the strongest baseline across multiple domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.