acceptodds
Under review as a conference paper at ICLR 2027

Learning Preferences of Strategic Users in Sequential Delegated Menu Design

Abstract

AI assistants often learn users' preferences from their choices and use this feedback to shape future recommendations. Strategic users, however, may choose to influence what the assistant offers next. We study this feedback loop in sequential delegated menu design, where an assistant repeatedly reveals options, learns from user choices, and maximizes its own payoff while protecting user welfare. We show that treating strategic choices as truthful can lead to learning incorrect preferences, whereas accounting for strategic behavior enables learning when feedback sufficiently distinguishes possible preferences. Yet, equilibrium feedback need not be informative: there can be an equilibrium in which all users provide the same feedback, leaving the assistant with only its best value without feedback. Moreover, when the assistant's behavior determines a terminal item for each possible preference profile, strategic feedback cannot outperform truthful feedback, even if it reveals the user's preferences. This conclusion has a sharp boundary: strategic feedback can outperform truthful feedback if the assistant observes feedback before randomizing future recommendations. If it randomizes before the user responds, we characterize exactly when truthful and strategic feedback have the same value. Experiments on synthetic interactions and preferences estimated from real-world datasets support these learning and value results.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.