Human-Like Generalization of Risk Preferences in LLMs
Abstract
The success of real-world interactions between humans and LLM-based chatbots depends on the character, preferences, and values of the LLM's "Assistant'" persona. However, the extent to which isolated measures of LLM behavior—e.g., evaluations using standard instruments from economics, cognitive science, or psychology—are predictive of realistic LLM chatbot behavior is unclear. In this work, we investigate this question of generalizability in the specific domain of *risk preferences*. We measure the risk preferences of 54 LLMs using three kinds of evaluations: surveys, lottery choice questions, and responses to human-written risky dilemmas. The first two evaluation categories are standard in psychology and experimental economics, and the last evaluation category reflects how LLM chatbots are used in practice. We find that survey evaluations predict the degree of riskiness of an LLM's responses to human-written risky dilemmas, while lottery evaluations do not, matching human risk preference generalization behavior. Our risky dilemma evaluation is conducted via a large-scale human experiment, in which participants write risk-related dilemmas from their own life and rate which LLM responses they prefer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.