RoleDrive: Scene Representations with Distinct Planning Roles for End-to-End Autonomous Driving
Abstract
End-to-end autonomous driving offers a promising route to safe and efficient navigation by learning to predict driving trajectories directly from sensor observations. However, many existing end-to-end planners mix different types of planning-relevant information within a unified representation, without explicitly differentiating their roles in trajectory planning, potentially limiting planning performance. To address this problem, we propose RoleDrive, an intermediate scene representation framework that explicitly organizes scene information into three complementary representations with distinct planning roles. Specifically, RoleDrive constructs Context Tokens for global scene context modeling and Map Tokens for fine-grained local spatial reasoning from shared visual features, while Evolution Tokens capture recent scene changes through residuals between aligned historical and current Map representations. These complementary representations are then progressively integrated through hierarchical trajectory reasoning. Without using any additional training data and with only 36.9M parameters, RoleDrive achieves a PDMS of 94.9 on NAVSIM-v1, surpassing the reported human-driver baseline, while also attaining state-of-the-art performance on NAVSIM-v2 and the photorealistic closed-loop HUGSIM benchmark. Code will be publicly released upon acceptance at https://github.com/K1ngAO/RoleDrive
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.