Learning When to Change Your Mind: SF-DPO for Multi-Turn Language Models
Abstract
Large language models (LLMs) can become trapped by their own mistakes in multi-turn interactions: once an erroneous answer enters the conversation history, it may increase the likelihood of subsequent errors. We refer to this history-dependent propagation as *error self-excitation*. However, suppressing such error propagation alone is insufficient for a robust multi-turn interaction, since models should both recover from erroneous answers while preserving correct ones under adversarial pressure and condition the revision on the validity of newly supplied information. We therefore formulate *multi-turn robustness* as a *state-dependent update problem* and introduce StateFork Direct Preference Optimization (SF-DPO), a lightweight post-training framework that learns when prior answers should be revised and when they should be preserved. SF-DPO holds the conversation history and current user input fixed and forks this state into preferred and rejected continuations through four complementary preference arms targeting *recovery*, *stability*, *valid-evidence updating*, and *fabricated-evidence resistance*. Finally, these state-forked preferences are optimized with DPO and chosen-response anchoring. Across three benchmarks and 33 model configurations from seven model groups spanning approximately 0.5B to 32B parameters, SF-DPO substantially improves *multi-turn robustness* across both *non-reasoning* and *inference-time reasoning* models while reducing *causal error propagation*, maintaining *evidence selectivity*, and requiring low *inference cost*.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.