MoLan: Motion-Language Understanding and Generation for Traffic Simulation
Abstract
Realistic, reactive, and controllable multi-agent traffic simulation is essential for the scalable development and evaluation of autonomous driving systems. Natural language provides an intuitive interface both for controlling how traffic scenes unfold and for querying the behaviors and interactions represented in them. Existing approaches, however, typically specialize in either language-conditioned motion generation or motion-grounded scene understanding. We introduce MoLan, a Motion-Language model that connects pretrained traffic and language towers through lightweight bidirectional bridge blocks. The resulting architecture supports native motion simulation and text generation, language-controlled closed-loop multi-agent rollouts, and motion-grounded traffic-scene understanding within a single model. MoLan outperforms task-specific baselines on both language-controlled simulation and motion-grounded question answering, while joint training across both directions further improves closed-loop realism, indicating that motion-understanding supervision can strengthen language-controlled generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.