Ego-Centric Joint Supervision for Planning-Oriented Autonomous Driving
Abstract
End-to-end (E2E) planners predict the future trajectories of the ego and the surrounding agents, yet recent methods supervise each agent marginally against its own trajectory. However, scenes in our training data and real-world driving data often require the planner to consider ego–agent interaction, which a marginal objective does not supervise explicitly. To address this, we propose CastAD, which trains a planning-oriented E2E planner under ego-centric joint supervision. Unlike conventional joint forecasting objectives, which treat all actors symmetrically, this supervision centers the scene on the ego, whose future is a decision in planning. Only the ego and its interaction-relevant agents, those with any predicted future that overlaps the spatial trajectory the ego followed, are supervised as one scene under the mode closest to the ego's spatial trajectory. A policy optimization stage then refines the stochastic ego policy on the joint-mode ego embeddings, which this supervision trains to encode ego–agent interaction, with rewards scored against map context and logged agent futures. CastAD reaches state-of-the-art results on closed-loop Bench2Drive (Driving Score 87.53, Success Rate 71.82) and on NAVSIM (PDMS 91.1). Video demos are available at https://castad-anon.github.io.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.