User Alignment Requires More Than Following Feedback
Abstract
Aligning language models from user interactions offers a direct path to personal assistants and agents that continually adapt to their users. Existing methods generally treat actionable feedback as ground-truth supervision: once a follow-up indicates a direction for improvement, the model is optimized toward the behavior inferred from it. This assumption is ill-suited to real user feedback, which is often incomplete, noisy, or open to revision. Such updates can prematurely commit the policy to weak evidence or unresolved choices that later feedback may undo. Therefore, learning from such naturally occurring feedback requires answering two core unresolved questions: \emph{how strongly should a feedback turn influence the policy, and which policy changes does it support? To address the first gap, we formalize it through feedback sufficiency, the extent to which a change implied by feedback is supported as a policy update. We then introduce Hindsight of Hindsight (HOH), which assesses this support by examining how the original user responds after that change is carried out. These retrospective assessments supervise an online sufficiency model that operates when feedback arrives. To address the second question, we propose User-Centric Multi-Teacher On-Policy Distillation (UMOPD), which constructs multiple feedback-compatible revisions rather than distilling a single feedback-conditioned teacher. UMOPD compares the policy shifts these revisions induce, reinforcing changes they share while preserving alternatives where they disagree. Predicted sufficiency then scales the resulting policy update. We validate UMOPD across three interaction scenarios: learning a user's preference, adapting when it changes, and integrating multiple preferences introduced over time. Under a naturalistic user simulator, UMOPD exceeds the strongest baseline in final preference win rate by points when learning an initial preference and by points after it reverses. When three preferences are introduced sequentially, UMOPD is the only method to exceed on each. Our code is available at https://anonymous.4open.science/r/umopd_hoh-2864.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.