Robust LLM-based User Simulation with Trained Latent-state Dynamics
Abstract
User simulators for training and evaluating LLM-based conversational agents must perform well off-policy (i.e., when facing novel agents). Recent work has argued that explicitly modeling user _latent state_ can improve simulator fidelity and off-policy robustness, but treating that state as a latent variable in LLM-based user simulation faces challenges such as intractability and unidentifiability given observational data. Consequently, latent-state simulation has generally been approached heuristically. We develop __RUSTLeD__, a principled probabilistic framework for latent-state learning that formalizes fine-tuning on latent-state labels as posterior-regularized maximization of the marginal likelihood of observed conversational data, smoothly interpolating between teacher distillation and autoencoding. Our _self-normalized importance sampling_ EM algorithm (SNIS-EM) addresses the problems facing current heuristic approaches. Beyond offering sound probabilistic foundations for latent-state simulation, __RUSTLeD__ yields simulators that generalize better to unseen agents than natural baselines, both in the likelihood of held-out human conversations and in reproducing human behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.