What you ask is what you get: Evaluating Robustness of Eliciting LLM Moral Profiles
Abstract
Understanding the mechanisms underlying moral reasoning is central to designing language technologies aligned with human values and preferences. Despite growing research on this topic, existing methods that systematically elicit the moral profile from Large Language Models (LLMs) through surveys show lack of reliability and generalizability to downstream behaviors. In this work we address this gap by presenting a methodological framework that combines psychological questionnaires and open-ended interviews to elicit differences in moral reasoning between LLMs, and test whether these differences shape model behavior in a downstream task. We apply this approach to two established psychological theories - Moral Foundations Theory and the Theory of Basic Values — across six open-weight models. Our results show that changing the elicitation method from psychological questionnaires to open-ended interviews leads to the elicitation of different moral profiles. The profiles of LLMs differ between each other both when they derived from closed-ended replies to questionnaires (Kruskal-Wallis ≈0.46–0.49) and from open-ended interviews (≈0.11–0.14), though with differences in magnitude. Steering models with their own elicited moral beliefs reduces inter-model agreement on hate speech classification, showing that our framework enables quantifying the effects of different moral profiles on a downstream classification task. Thus, our study presents a first step towards a moral profile elicitation method that is explanatory of downstream behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.