Learning User Representations to Improve Simulation Fidelity
Abstract
User simulators enable interactive training and evaluation of AI assistants at scale. Yet current simulators do not reproduce the diverse behaviors of real users very well, leading to assistants that struggle to generalize to underrepresented users. We identify user representations as an overlooked bottleneck in user simulation. These are natural-language descriptions that simulators are conditioned on to reproduce a particular user's behavior. A representation's quality becomes apparent in hindsight, through the behavior it induces in a simulator. To evaluate and optimize user representations, we formulate a conversation reconstruction task: given an original conversation, we generate a user representation, condition a simulator on it, simulate a conversation, then measure fidelity by comparing user behavior in the original and simulated conversations. This comparison also gives hindsight feedback, which identifies where the representation falls short and how it can be improved. We introduce hindsight distillation, an online training procedure that distills feedback-conditioned improvements into the representation generator, enabling it to generate improved representations directly from the original conversation. Across 700 conversations spanning seven datasets, hindsight feedback improves user representations at test time, and hindsight distillation yields UserRepGen-4B, whose generated representations enable more faithful simulation than those generated by GPT-5.5. Beyond conversation reconstruction, the learned representations also improve fidelity for unseen tasks, better preserve the population-level diversity of real users, and increase user re-identification accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.