Same Choice, Different Response: Evaluating LLM Simulators Under Attribute Changes
Abstract
Large language models (LLMs) are increasingly used to simulate human choices, but reproducing observed choices may not predict how people respond when decision attributes change. We study interventional response alignment: agreement between human and model changes in the full choice-probability vector. We show that low observational choice risk does not, without additional assumptions, guarantee low response risk. On SwissMetro, we evaluate responses to travel-time and fare changes against fitted population-level human references. Prompted LLMs show response discrepancies even when their baseline choices match recorded human choices. More strikingly, on a fixed respondent-disjoint test set, choice-only SFT and DPO raise a 7B model's accuracy from 44.3% to 69.3–70.8%, yet improve time alignment while worsening fare alignment and the prespecified combined endpoint. Supervising paired response differences reduces the combined gap by 41% relative to SFT and outperforms choice augmentation on the same inputs. This gain does not extend to a held-out attribute, comes with worse negative log-likelihood, and has an unresolved accuracy difference. Because responses to the constructed changes are unobserved, these findings concern agreement with fitted population references rather than individual counterfactual fidelity. They show why choice accuracy alone is insufficient to evaluate simulators intended for use under change.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.