MindWorld: Iterative Co-Optimization of Mental World Models and Policies for Social Intelligence
Abstract
Language agents increasingly operate in interactive environments, yet their understanding of users and counterparts is often not maintained as an explicit, evolving component of the decision process, limiting their ability to adapt as interactions evolve. We introduce MindWorld, a framework that explicitly couples a Mental World Model with a Policy Model. The Mental World Model dynamically tracks beliefs, desires, and intentions from the evolving interaction history, while the Policy Model conditions on these mental states together with the task context to generate actions. To improve long-horizon credit assignment, MindWorld combines trajectory-level outcome feedback with dedicated step-level rewards and converts them into step-wise relative advantages for reinforcement learning. We further introduce an iterative co-optimization procedure. The two models are trained in alternating phases using a trajectory-level outcome reward and step-level rewards from fixed scorers. Experiments on SOTOPIA, SOTOPIA-Hard, ALFWorld, and WebShop demonstrate consistent gains across social, household, and web-based interaction. With Qwen3-8B under self-play, MindWorld obtains GOAL/OVERALL scores of 9.27/4.03 on SOTOPIA and 8.16/3.94 on SOTOPIA-Hard. On ALFWorld and WebShop, the reported results exceed those of the evaluated general RL baselines, and integration with task-specific methods yields additional gains. These results show that explicitly modeling evolving mental states, together with fine-grained credit assignment and iterative policy optimization, provides an effective and general approach for improving interactive language agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.