Does Behavioral Fidelity Predict Reliable Policy Selection? Decision-Anchored Validation of LLM User Simulators
Abstract
Large language models (LLMs) are increasingly used to simulate user responses in recommender systems, supporting agent training and policy evaluation. Existing studies primarily assess the behavioral similarity between simulated and real users, yet it remains unclear whether such fidelity reliably reflects the quality of recommendation policy selection. To address this question, we characterize LLM-based user simulators from three complementary perspectives: distributional fidelity, conditional response fidelity, and treatment-response fidelity. We further construct simulator conditions with systematically different fidelity profiles through probability calibration, behavior cloning, treatment-aware training, and controlled information perturbations. Using randomized real-user logs, we estimate reference values for supported recommendation policies and directly compare simulator fidelity with policy ranking, pairwise decision correctness, and decision regret. Experiments and cross-domain stress tests show that making a simulator more behaviorally similar to real users does not necessarily make it more reliable for comparing or selecting recommendation policies. The relevant code can be obtained at https://anonymous.4open.science/r/SimFidelity-523F.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.