Learning from What Users Do Not Say: Self-Evolving Assistants by Internalizing Latent Thoughts
Abstract
Large language models (LLMs) are increasingly used as interactive assistants, yet their ability to understand users remains limited by partial observability: user utterances often reveal only incomplete information about the underlying intents, preferences, and concerns. In this paper, we study whether such latent user information can serve as privileged supervision during training and be internalized into the assistant for real-world interactions. We propose OASIS, a three-stage framework that learns from latent user thoughts without relying on a stronger teacher model. OASIS first uses user simulators to expose latent user states for supervised cold-start training, and then performs on-policy asymmetric self-distillation, where the same assistant learns from the information gap between a privileged view with latent thoughts and an observable view with dialogue only. Experiments on open-ended general assistance and goal-oriented commercial customer service show that OASIS consistently improves assistant performance across model scales and evaluation settings, with on-policy self-distillation providing up to 52.4% and 22% further improvements over the SFT initialization on WildBench and DianJin-CSC, respectively. These results demonstrate that latent user information can be effectively internalized into the assistant through asymmetric self-distillation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.