acceptodds
Under review as a conference paper at ICLR 2027

Learn the User, Serve Them Better: From User Models to Personalized LLM Decisions

Abstract

An online personalized LLM learns from user feedback. While pairwise questions reveal which response a user prefers, absolute acceptability questions ask whether a response is good enough. With a limited feedback budget, our proposed DUAL (Dual-feedback User-Adaptive Learning) learns each user's preferences and acceptance threshold to decide whether to serve either response or abstain. An established information-gain score selects the question type, while a calibrated cutoff and quota completion determine when to ask. Under independent uniform candidate features, a correctly specified logistic user with moderate true logits, and continuing answer quotas, we prove that DUAL's original recursive estimates converge almost surely to the true preference weights and acceptance threshold. We also characterize information coverage and bound the online estimate's distance to a joint fit. A controlled development simulation isolates question choice: at identical question times and 940 answers, information selection achieves 15.85% area under the serving-mistake learning curve (AUC), versus 16.45% for a separately calibrated fixed allocation. A common-calibration development study compares complete policies at 940 answers: DUAL achieves 14.735% AUC versus 15.206% for adaptive acceptability-only questioning. These simulations use PersonalLLM features and controlled acceptability feedback. Exploratory PRISM results leave the benefit of mixing feedback after population initialization unresolved. Community Alignment supports personal choice-probability prediction and information-based question selection over entropy. The complete policy's human serving benefit remains open.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.