Unsupervised Preference Alignment of LLMs from Multi-Turn Dialog Logs
Abstract
Multi-turn preference alignment is hindered by long-range context dependencies and costly human annotation. We present a fully unsupervised framework that constructs Direct Preference Optimization (DPO) data from production dialog logs without human labels or LLM judges during pair construction. Responses from a stronger deployed LLM system are chosen; weaker-model responses to the same context are rejected. This cross-model construction creates a source-model confound: preference-alignment methods, including DPO, can exploit *generative fingerprints*—model-specific stylistic and distributional patterns—instead of semantic quality. We formalize the confound and show that shared, sentence-level paraphrasing of both responses can reduce fingerprint separability while bounding semantic-signal drift; our pipeline combines normalization, length balancing, and both-side paraphrasing. Across 77K–11M pairs from 12M real multi-turn conversations, the regularized pipeline improves with scale, whereas standard DPO overfits at all scales. Phi-4 (14B) gains up to +7.8pp held out, and +6.2pp, +3.1pp, and +5.0pp on WildChat, AlpacaEval, and Arena-Hard prompts under our binary rubric. Using Phi-family preference data, Qwen-2.5-14B gains +6.7pp held out and +2.9pp on WildChat. Two additional judge families and a target-independent paraphraser control confirm the key held-out gain, and public WildChat data replicates the pipeline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.