Isolated Consistency: Improve Multi-Turn Reasoning via Turn-Level Isolation
Abstract
Multi-turn conversation lets users specify a task gradually, but language models often fail even after all requirements have been provided. One difficulty is that the model starts answering before it has the full task. These early responses remain in the conversation and can shape later answers, even when new requirements invalidate their assumptions. Restating the complete task clarifies what the model should solve, but does not tell it which earlier responses to trust. These responses may offer useful intermediate reasoning while also carrying mistaken assumptions, so neither keeping nor discarding them all is an obvious choice. We observe that responses containing errors often make up only a minority of assistant turns. This suggests that isolating earlier responses can keep these errors out of most candidate contexts, rather than exposing every attempt to them. Instead of training a selector to decide what to retain, we introduce Isolated Consistency (IC), a training-free method compatible with both white-box and black-box models that revisits the task with each earlier response in isolation and then combines the resulting answers through voting. Concretely, IC generates one candidate answer for each earlier assistant turn. Each candidate uses all user messages and the supplied clean recap, but only the assistant response from that turn, with all other responses masked. IC then groups equivalent answers and selects the most common one. Across seven models and four analytical tasks, IC improves recovery beyond the strongest evaluated multi-turn baseline, raising macro-average accuracy by 3.19 percentage points with the same number of candidates and voting rule. These gains extend even to GPT-5.5, where IC improves accuracy by 25.87 points over standard multi-turn prompting and by 3.54 points over the strongest baseline, showing that stronger models still benefit from reconsidering which earlier responses to use.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.