SiWM: Learning Planning-Relevant World States via Semantic Integration
Abstract
End-to-end autonomous driving often learns planning representations through predefined perception tasks such as object detection, lane structure estimation, and motion prediction. However, it is difficult for a finite set of tasks and labels to cover all scenes and state changes that affect driving decisions on open roads, and this supervision typically requires extensive manual annotation. We quantify perception task losses' impact on closed-loop driving and investigate how to learn environment representations that support it without these losses. We propose the Semantics-Integrated World Model (SiWM), which integrates visual features and continuous planning semantics from a large language model (LLM) into a shared world state, so that both types of information contribute jointly to state learning and recurrent updates. Through action-conditioned future prediction and feature reconstruction, the model learns how the environment evolves and preserves scene information. It generates ego trajectories from the fused state, incorporating driving task information throughout state learning and planning. After 4 task-specific training epochs, SiWM outperforms the direct planning baseline in closed-loop driving. The gains are 20.43 points in Driving Score (DS) and 21.36 percentage points in Success Rate (SR). In some scenarios involving sudden changes, it also achieves better closed-loop performance than ORION, which uses explicit perception task supervision. Further experiments show that integrating planning semantics into world model learning outperforms fusing planning semantics only at the planner, while multi-token representations substantially alleviate collapse from converging state directions across scenes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.