Identifying Reward Factors for Sparse Preference Elicitation
Abstract
Personalized reward models represent each user's reward as a weighted combination of shared factors. Predictive accuracy alone does not give these factors a stable meaning: an invertible recombination of the factors, offset by a corresponding change in the weights, preserves every prediction while changing the meaning of each weight. We study which factors a population's comparison data identify and how many comparisons are needed to recover a new user's sparse composition once the dictionary is fixed. When each reward is a positive scale times a simplex composition, a population with a specialist for every factor generates a reward cone whose extreme rays are exactly the factors. The dictionary and user compositions are then identifiable up to permutation within the anchored model class. Without anchors, the extreme rays correspond to extreme population mixtures and may exclude sparser newcomers. For a fixed dictionary, the simplex constraint makes the sign of a realizable group contrast test the user's total weight on that group. The decision is independent of the user's scale, although its reliability depends on that scale. CS-Elicit recovers an -sparse support by testing tree nodes, using noisy comparisons at answer bias . This count matches a counting lower bound up to logarithmic factors, and the weight estimation rate is independent of . Across language and control tasks, extraction improves dictionary stability across independent data draws. On PersonalLLM, extracting factors from LoRe utilities raises the cross-draw dictionary correlation from to . In the E4 scaling experiment, CS-Elicit reaches held-out accuracy for -sparse users with an extracted dictionary in questions, compared with for mutual information selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.