Structuring Privileged States with LLMs for Robust World Models
Abstract
Visual world models enable control directly from high-dimensional visual observations through latent imagination, yet their learned representations can remain brittle under distribution shifts. While prior work studies generalization under complex visual distractions, we discover that simple changes to environment configurations and physical dynamics are sufficient to cause substantial performance degradation under zero-shot evaluation. Privileged states expose rich structural information about the underlying environment and have been broadly adopted to support robust generalization, but directly predicting raw states may require modeling task-irrelevant or visually inaccessible quantities, while manually designed auxiliary targets scale poorly across diverse tasks. In this work, we propose a framework that harnesses the reasoning capabilities and commonsense priors of Large Language Models (LLMs) to automatically structure privileged states into compact, task-relevant semantic properties. These LLM-derived properties serve as auxiliary targets that explicitly regularize the world model toward representations that remain stable across distribution shifts. To seamlessly integrate this supervision, we introduce an adaptive multi-loss calibration mechanism to dynamically balance the semantic prediction objective against the original world model losses, removing the need for task-specific manual weight tuning. Across Meta-World, DeepMind Control Suite, and CARLA, our approach consistently achieves substantial zero-shot generalization improvements over strong model-free and model-based baselines under both held-out configuration and physical dynamics shifts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.