Beyond the Preference Monolith: What Predicts Human Taste?
Abstract
In subjective, open-ended scenarios, individuals and even whole communities may develop implicit notions of quality that are unarticulated. What lies behind these quality judgments? Past work has implicitly modeled these judgments by training reward models on datasets of human preferences, or constructing LLM-as-a-judge methods. Instead, in this work, we seek to explain the underlying factors that predict human quality judgments. We propose Atomic Metrics, an interpretable framework that constructs fine-grained explanatory criteria across a population of pairwise-preferences. This method first induces atomic criteria by contrasting preferred and dispreferred responses, refines them using preference feedback, and learns the relative importance of each criteria. We use Atomic Metrics to improve LLM judges, showing that the learned Atomic Metrics are highly correlated with human preferences across a range of tasks and demographic groups. We also show that these preferences can be highly contextual and distinct, depending on the task context, and the community whose preferences are being studied. We further demonstrate that Atomic Metrics can be used to improve the performance of models, by guiding their generations in aspects that humans have revealed are important to them. Finally, we demonstrate the potential benefits of Atomic Metrics for individual personalization, where it is possible to learn specific tastes of individuals from just a few examples. Thus, this study seeks to shed light on the latent factors that predict observed human preferences in subjective open-ended tasks, as well as improve the training, evaluation, and personalization of large language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.