acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Change Your Mind: SF-DPO for Multi-Turn Language Models

Abstract

Large language models (LLMs) can become trapped by their own mistakes in multi-turn interactions: once an erroneous answer enters the conversation history, it may increase the likelihood of subsequent errors. We refer to this history-dependent propagation as *error self-excitation*. However, suppressing such error propagation alone is insufficient for a robust multi-turn interaction, since models should both recover from erroneous answers while preserving correct ones under adversarial pressure and condition the revision on the validity of newly supplied information. We therefore formulate *multi-turn robustness* as a *state-dependent update problem* and introduce StateFork Direct Preference Optimization (SF-DPO), a lightweight post-training framework that learns when prior answers should be revised and when they should be preserved. SF-DPO holds the conversation history and current user input fixed and forks this state into preferred and rejected continuations through four complementary preference arms targeting *recovery*, *stability*, *valid-evidence updating*, and *fabricated-evidence resistance*. Finally, these state-forked preferences are optimized with DPO and chosen-response anchoring. Across three benchmarks and 33 model configurations from seven model groups spanning approximately 0.5B to 32B parameters, SF-DPO substantially improves *multi-turn robustness* across both *non-reasoning* and *inference-time reasoning* models while reducing *causal error propagation*, maintaining *evidence selectivity*, and requiring low *inference cost*.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.