CoWISE: Co-Evolving World Models with Policies through Semantic–Behavioral Coordination
Abstract
World models enable language agents to predict the consequences of their actions and improve their policies through simulated interaction. To train these models, existing approaches typically use trajectories collected from agent-environment interactions to reproduce real observations. However, even semantically similar simulated observations can still induce different policy actions (e.g., simulated observations closely match the real observation in overall content but miss key objects, leading the policy to take incorrect actions), which we identify as behavioral mismatch. This behavioral mismatch grows as the policy continues learning while the world model remains fixed, thereby further misleading the policy into choosing wrong actions and degrading its performance. To reduce this mismatch, we propose **CoWISE**, the first dual-level co-evolution framework that continually adapts the world model to the evolving policy through semantic-behavioral coordination. Specifically, at the policy–world level, CoWISE enables the world model and the policy to co-evolve by leveraging each other’s predictions. At the semantic–behavioral level, it adaptively balances real-world observational supervision and policy-derived behavioral supervision. To evaluate CoWISE, we conduct extensive experiments across 8 benchmarks, demonstrating that CoWISE improves policy performance and learns a world model with higher behavioral fidelity. Furthermore, using this learned world model for test-time scaling yields further gains, outperforming combinations of separately trained policy and world model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.