acceptodds
Under review as a conference paper at ICLR 2027

Contrastive Self-Distillation from User Feedback

Abstract

Large language models deployed in production generate a large volume of real interactions with users. Users rarely provide explicit scores for responses, but they often correct errors, clarify requirements, or express dissatisfaction through follow-up messages, making subsequent user messages a cheap yet rich learning signal. A class of self-distillation methods, such as SDPO, uses the next user message as hindsight evidence. After reading it, the model re-scores the preceding response, and the resulting token-level score change is distilled back into the model. Although distilling this score change can yield strong empirical gains, we show that the change itself contains a systematic confound: even empty templates or unrelated messages induce sizable, response-dependent score shifts. Directly distilling the full score change therefore writes this component into the update as well, which is unrelated to the current feedback. Theoretically, we show that this component is independent of the message and enters the score additively, so it cancels when the scores under the real message are contrasted with those under randomly sampled reference messages, without having to estimate it. Building on this result, we propose CoSD (Contrastive Self-Distillation), which distills how much more the real message shifts the model's scores than a few reference messages sampled from other interactions, requiring no external reward model, preference annotations, or additional generation. Furthermore, CoSD allows cross-session user history to inform how the current message is interpreted. Since message-independent effects of the shared context cancel in the contrast, all branches can condition on the user's history without distilling the user-dependent normalization it introduces. Extensive experiments across multiple backbone models, two real-world interaction corpora, and both on- and off-policy training show that CoSD improves over the initial model and outperforms SDPO.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.