Do LLMs Know How a Person Handles Everyday Self-Control Conflicts? A Person–Episode Test with Experience Sampling
Abstract
Self-control in daily life is a process: when people face a conflict, they select strategies that depend on who they are and on the situation, appraise how well they managed, and monitor and evaluate their strategy use. Large language models (LLMs) match human performance on several social-inference tests and are increasingly used to model and support people, yet their evaluations score outcomes, such as keyed answers, group-level rates, and questionnaire scores, leaving open whether a model reproduces the process by which a specific person handles a specific episode. We test 22 LLMs, prompted in English and German with increasingly detailed person descriptions, on 3,725 experience-sampled self-control conflicts from 468 participants, and score each predicted step (the person's strategies, success, and metacognition) against the person's own reports and reference predictors. At the selection step, the models list more strategies than people report, predict similar sets for different people, and recover the reported strategies less well than a rule that lists the same number of the most frequent strategies. At the appraisal step, person information improves their ranking of people by typical success and changes their tracking of a person's episodes only slightly. At the metacognitive step, their ratings follow reported success and strategy counts, ranking metacognitive knowledge weakly and missing regulation, which the conflict histories carry. The models thus capture what people generally do and how people differ, and miss how a person's situation shapes strategy choice, how strategies shape the appraisal of an episode, and how strategy use across episodes reflects metacognition. Because each step looks plausible on its own, evaluations of LLMs as models of individuals should test this process directly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.