acceptodds
Under review as a conference paper at ICLR 2027

Predicting the Critic: Where Feedback Goes Decides What It Teaches

Abstract

Emergent misalignment (EM) is usually studied on datasets that end with the harmful assistant answer, whereas in real conversations the user often reacts to it. We fine-tune five families of open-weight models on identical question and answer pairs to which such a reaction is added, in front of the question as context or after the answer with or without loss. Critical reactions in front of the question reduce EM in all five families, but adding the note back to the prompt at evaluation time restores EM to at least the baseline level, i.e. the reduction depends on the context. Included in the loss, the reaction mainly changes the model's predictions of user reactions (in all five families), and EM is not consistently reduced, with small effects of different sign across families, reaction types, and domains. If reactions are correlated with answer quality, the model's predicted reactions are too and can be used to re-rank its own samples, with almost all measured EM removed in a symmetric sweep over feedback reliability by choosing the least-criticized of 16 samples at a reliability of 70% or higher. When reactions are random, the same choice can increase EM. These results suggest that for post-training on conversational data, the choice of loss masking and the reliability of the feedback deserve attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.