Future-Grounded Supervision: Improving LLM Intelligence with Real-World User Behavior Data
Abstract
Large language models are increasingly deployed in real-world applications, from personal assistants that take actions based on users' long-term behavior histories to recommendation and advertising systems that predict what users will do. These applications demand the user dimension of LLM intelligence: the ability to infer a person's preferences, intent, and likely future actions from weeks to years of user behavior data. Yet current models remain limited in this user dimension of intelligence: they are trained on static text or synthetic data, neither of which captures how user behavior evolves over time. Real-world user behavior data, by contrast, is already available at scale on content platforms and in personal-assistant agents. These records capture sequences of user actions, making the user's actual next action a natural supervisory signal with the potential to improve LLM intelligence. To this end, we propose *future-grounded supervision*, a paradigm in which the training signal for LLMs comes from the user's real future behavior. We further introduce **Vela**, a framework that operationalizes this paradigm by converting real user behavior data into trainable tasks and trajectories. From the de-identified behavior records of 100,000 users, Vela generates 100,000 tasks across 91 real-world task templates. Difficulty-based filtering and trajectory acceptance criteria jointly retain approximately 75,000 tasks with selected training trajectories. Fine-tuning Qwen3.6-35B-A3B on these training trajectories improves performance over the base model by up to 11.5 points on user-related tasks and by an average of 4 points on general benchmarks. Scaling experiments further show that training on more user behavior data improves user-related performance. Reinforcement learning after fine-tuning outperforms reinforcement learning from the base model by up to 7.7 points on user-related tasks. These gains suggest that future-grounded supervision on real-world user behavior data can improve not only user-centered capabilities but also the broader intelligence of LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.