CARLAGen4D : A Multi-Agent Framework for Text-Driven 4D Scene Generation in CARLA
Abstract
Autonomous driving research heavily relies on high-fidelity simulators such as CARLA, where evaluation is almost exclusively performed on a limited set of predefined, artist-designed towns. Prior work in simulation-based evaluation largely depends on these fixed maps, constraining scenario diversity and limiting the ability to test generalization to novel environments and edge cases. As a result, current simulation-based evaluation pipelines struggle to scale with the complexity and variability required for robust autonomous driving validation. To address this gap, we present an end-to-end framework for text-to-4D simulation generation that directly enables controllable, scenario-driven evaluation of AD systems. Given a natural language prompt describing a testing scenario (e.g., unusual road topology, adversarial traffic behavior), our multi-agent system generates a complete, temporally consistent simulation, jointly synthesizing road topology, background assets, and dynamic agent behaviors from scratch. This removes reliance on fixed maps and enables on-demand creation of targeted evaluation scenarios. Since direct map generation is prohibitively expensive in token space, we design a domain-specific language that enables efficient map synthesis. We further benchmark driving policies across diverse generated road layouts and traffic scenarios, revealing their sensitivity to map configuration and scenario diversity. Beyond evaluation, the generated scenes support a range of downstream applications, including RGB-to-BEV perception, scene narration and captioning, and transfer to real-world imagery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.